Fireworks Ember-1 as a Cursor and Cline BYOK Backend in 2026: Kimi K3 Scores at 40% Less Token Spend — and a Two-Week Catch

fireworksember-1kimicursorclinebyokapicostpricing

TL;DR: Ember-1, released September 24, 2026, is Moonshot’s open-weight Kimi K3 post-trained by Fireworks Research to reason in fewer tokens. It costs exactly what K3 costs on Fireworks — $3 input / $15 output per million — but Fireworks’ production A/B on real coding workloads measured 71.3% fewer reasoning tokens and 39% lower total token spend for near-identical benchmark scores. The catch: it’s a research preview with a two-week serverless window, and nobody outside Fireworks has replicated the numbers yet.

Ember-1 (Fireworks)Kimi K3 (Moonshot first-party)Gemini 3.8 Flash
Best forK3-quality agent runs at a real discountFrontend generation where K3 leads outrightVolume loops on the tightest budget
Input / Output per 1M$3.00 / $15.00 ($0.30 cache read)$3.00 / $15.00 ($0.30 cache hit)$0.75 / $3.75 through Dec 31 → $1.50 / $7.50
SWE-bench Verified92.2% (Fireworks-run)93.2% at max effort (Fireworks-run)not published; 61.6% SWE-Bench Pro (Google-run)
The catchResearch preview, two-week serverless window~2x output verbosity runs the $15/M meter hotPrice doubles January 1, 2027
Context window1,048,576 tokens1M1M in / 65,536 out

Honest take: If Kimi K3 is already your Cline backend, point the model string at Ember-1 today — same sticker, roughly 40% less metered output for the same work, per Fireworks’ own A/B. But treat it as a two-week trial, not a migration: keep your K3 config one toggle away until Fireworks commits to serving Ember-1 permanently and someone independent reruns the benchmarks.

What did Fireworks ship on September 24?

Ember-1 is the first release from Fireworks Research, and it’s a post-train, not a new model: Fireworks took Moonshot’s open-weight Kimi K3 — the 2.8-trillion-parameter model we covered as a Cursor and Cline backend in July — and fine-tuned it on Fireworks’ own training stack to reach the same answers with shorter reasoning traces. The company says it ran more than 50 training experiments and over 200 evaluations to get there.

The spec sheet, verified against Fireworks’ model page and the OpenRouter listing on September 28, 2026: model ID accounts/fireworks/models/ember-1, a 1,048,576-token context window, image input, and function calling — the last one matters, because tool calls are what Cursor’s agent mode and Cline live on. It’s served on Fireworks serverless, and it’s already listed on OpenRouter as fireworks/ember-1 and on Vercel’s AI Gateway for anyone who’d rather not open another billing account.

Two words in the launch material deserve more attention than the benchmarks: research preview. Ember-1 ships with a two-week serverless window unless Fireworks makes it permanent. That’s not a footnote — it’s the difference between a backend you migrate to and one you evaluate. More on that in the “when not to use it” section.

Does “40% fewer reasoning tokens” hold up?

The claim is specific and, unusually for a launch post, comes with a production experiment attached. Fireworks says Ember-1 matches K3 with 35–50% shorter reasoning traces, and in live A/B tests on real coding workloads measured reasoning tokens down 71.3% and total token spend down 39%.

The benchmark table, all Fireworks-run, comparing Ember-1 against Kimi K3 at maximum reasoning effort:

BenchmarkEmber-1Kimi K3 (max effort)
SWE-bench Verified92.2%93.2%
Terminal-Bench 2.182.0%80.9%
DeepSWE 1.1wins—
SWE-Interactloses narrowly—

Read that as a wash on quality: Ember-1 wins outright on Terminal-Bench 2.1 and DeepSWE 1.1, loses by a point on SWE-bench Verified, and loses narrowly on SWE-Interact. A one-point trade on SWE-bench Verified for 39% less spend is a trade most people running agent loops should take.

Two caveats before you rebuild your budget around it. First, the token savings are wildly uneven by task type: 51.9% on Terminal-Bench against just 5.9% on τ²-Airline. Coding and terminal-agent work — the workloads this site cares about — sit at the favorable end, but if your sessions skew conversational, the discount shrinks toward nothing. Second, every number above was produced by Fireworks, the company selling the tokens. No independent replication existed as of September 28. We flagged the same thing when Google benchmarked its own Gemini 3.8 Flash; vendor-run numbers are directionally credible and provisional until someone neutral reruns them.

Why does this work at all? Kimi K3’s defining cost problem was never its sticker price — it was verbosity. Artificial Analysis measured K3 generating roughly twice the output tokens of comparable models, around 120K output tokens per agentic task. At $15 per million output tokens, always-on maximum reasoning is expensive. Ember-1 is a direct attack on that specific meter.

What does Ember-1 cost against the backends you’re already using?

Sticker price first: $3.00 input / $15.00 output / $0.30 cache read per million tokens — identical to what Kimi K3 costs on Fireworks. The pitch isn’t a cheaper rate; it’s fewer metered tokens for the same work.

On our standard session shape — 20K input tokens, 7K output, the same math used across our backend reviews — with all prices verified September 28, 2026:

BackendInput / Output per 1MPer sessionPer 1,000 sessions/mo
Muse Spark 1.3 (Contributor tier — Meta trains on your data)$0.10 / $0.20$0.003$3
Gemini 3.8 Flash (intro, through Dec 31)$0.75 / $3.75$0.041$41
Muse Spark 1.3 (standard)$1.25 / $4.25$0.055$55
Claude Sonnet 5 (~30% tokenizer overhead)$2.00 / $10.00$0.143$143
Ember-1 / Kimi K3 (sticker)$3.00 / $15.00$0.165$165
Claude Fable 5.1$10.00 / $50.00$0.550$550

The identical-shape comparison actually understates Ember-1’s position, because its whole point is that the same job emits fewer tokens. Apply Fireworks’ measured 39% total-spend reduction and that $165 per thousand sessions becomes roughly $100 — cheaper than Sonnet 5’s effective rate, for a model whose (vendor-run) SWE-bench Verified score sits in frontier territory. Against Fable 5.1, the queue math is starker: a Cline session that emits 50K output tokens costs $0.75 on Ember-1 and $2.50 on Fable 5.1, before Ember-1’s token reduction is counted.

The table also shows where Ember-1 doesn’t compete. Gemini 3.8 Flash at $41 per thousand sessions is a quarter of Ember-1’s sticker cost, and Muse Spark 1.3’s Contributor tier is two orders of magnitude cheaper if you’re willing to let Meta train on your prompts. Ember-1 is not a budget backend. It’s a mid-to-premium backend that claims premium-model output quality with a self-funding discount — the right comparison set is Sonnet 5 and Fable 5.1, not the Flash tier. The full market view is in our cost comparison across every tier.

How do you set up Ember-1 in Cursor?

Ember-1 is not in Cursor’s native model picker as of September 28, so this is a BYOK job through the OpenAI-compatible override:

  1. Create an API key at app.fireworks.ai/api-keys.
  2. In Settings → Cursor Settings → Models, open the API Keys section, paste the key under the OpenAI key field, and enable Override OpenAI Base URL with https://api.fireworks.ai/inference/v1.
  3. Add accounts/fireworks/models/ember-1 as a custom model name, then select it in Chat or Agent.

Before blaming either tool when something errors, prove the key and model string resolve with a 30-second sanity check:

curl -s https://api.fireworks.ai/inference/v1/chat/completions \
  -H "Authorization: Bearer $FIREWORKS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "accounts/fireworks/models/ember-1", "messages": [{"role": "user", "content": "Reply with exactly one word: ready"}]}'

A working setup returns JSON with "model" echoing the Ember-1 ID and your reply under choices[0].message.content, plus a usage block — get in the habit of reading it, since output-token counts are the entire value proposition here. A 401 with an error object means the key, not the model.

The standing Cursor BYOK caveats apply, verified against Cursor’s own API-keys documentation. Custom keys work in Chat and Agent only: Cursor Tab autocomplete, Auto model selection, Cloud and Background Agents, Automations, and the Cursor CLI all run on Cursor-routed models and ignore your key entirely. And the base-URL override is global — while it’s active, Cursor’s built-in models are unavailable, so flip it off when you want Sonnet back. Newer Cursor builds don’t show the override field on every plan; if yours lacks it, OpenRouter is the cleaner path — fireworks/ember-1 through OpenRouter’s standard endpoint, one account, no base-URL surgery.

How do you set up Ember-1 in Cline?

Cline is the easier home, because Cline has a first-party Fireworks provider — no OpenAI-compatible workaround needed:

  1. Open Cline’s settings and set the API provider to Fireworks AI.
  2. Paste your key from app.fireworks.ai/api-keys.
  3. Set the model ID to accounts/fireworks/models/ember-1.

If your Cline build’s Fireworks provider hasn’t caught up to the September 24 release, the OpenAI Compatible provider works identically: base URL https://api.fireworks.ai/inference/v1, same key, same model ID, context window 1,048,576, and function calling on. Cline’s agent loop is also where the token reduction pays most visibly — Fireworks’ biggest measured savings came on terminal-agent-shaped work, which is most of what Cline does all day.

One planning problem we hit immediately, and the fix: the two-week research-preview window means a config you set today could point at a retired endpoint in mid-October. The solution is to treat the K3 fallback as part of the setup, not an afterthought — keep a second Cline provider profile pointing at Kimi K3 (Moonshot first-party at api.moonshot.ai/v1, model kimi-k3, same $3/$15) so the fallback is a dropdown change rather than an outage. Since Ember-1 is post-trained from K3, behavior differences between primary and fallback are as small as fallbacks get.

When should you NOT put Ember-1 in your stack?

Four cases where the answer is something else:

  • You need a backend you can budget past mid-October. The two-week serverless window is real until Fireworks says otherwise. If you’re setting a team standard or wiring CI to a model endpoint, that’s disqualifying today — evaluate now, commit only if Fireworks makes it permanent.
  • Your work is accuracy-critical multi-file refactoring. Ember-1 gives up a point on SWE-bench Verified (92.2% vs 93.2%) and loses on SWE-Interact. A point is noise on most work; on the runs where you’d have picked Fable 5.1 anyway, the calculus in our Opus 5 default cost analysis still applies.
  • Your sessions aren’t coding-shaped. The token savings ranged from 51.9% down to 5.9% depending on task type. Conversational or retrieval-heavy workloads keep the $3/$15 sticker without the discount that justifies it — Gemini 3.8 Flash at a quarter the price serves those better.
  • You want frontend generation specifically. Kimi K3’s headline win was taking #1 on a human-preference frontend arena. Fireworks published no frontend-specific eval for Ember-1, so if UI generation is your workload, stay on the model with the measured result until someone tests the post-train against it.

And one non-case: waiting for the weights. Kimi K3’s weights are on Hugging Face under Moonshot’s Modified MIT license, but Fireworks announced no Ember-1 weights release in the launch materials — as of September 28 this is a hosted-only model. If zero-API-cost local inference is the goal, a 2.8T-parameter base was never going to fit in a home lab anyway; our sister site’s VRAM-based local model guide covers what your GPU can actually serve, and the open-weight coding-model landscape lives at aifoss.dev.

Verdict

Ember-1 is the most interesting kind of release: no new capability claim, just the same model made cheaper to run — and priced so the savings flow to you rather than to Fireworks’ margin. For Cline users already on Kimi K3, switching is a one-string change with a built-in fallback, and the downside is bounded by a two-week trial window. For Cursor users, it’s a functional but clunkier BYOK setup that disables the native picker while active. For everyone, the honest status is: vendor-verified, independently unverified, and temporary until announced otherwise. Run it on your own repo for the preview window, read the usage blocks, and let your own token counts — not Fireworks’ A/B — decide whether it stays. We’ll revisit when the preview window resolves or an independent benchmark run lands, whichever comes first.

FAQ

What’s the exact model ID and endpoint? accounts/fireworks/models/ember-1 on Fireworks’ OpenAI-compatible endpoint at https://api.fireworks.ai/inference/v1. On OpenRouter it’s fireworks/ember-1; it’s also on Vercel’s AI Gateway.

Is Ember-1 actually 40% cheaper than Kimi K3? Same sticker price — $3/$15 per million tokens on Fireworks. The savings come from emitting fewer tokens: Fireworks’ production A/B measured 39% lower total token spend on real coding workloads, with reasoning tokens down 71.3%. The reduction varies hard by task type (51.9% on Terminal-Bench, 5.9% on τ²-Airline), and no independent party has replicated it yet.

What happens after the two-week research preview? Fireworks hasn’t said. The launch terms give Ember-1 a two-week serverless window unless it’s made permanent. Keep a Kimi K3 provider profile configured as a fallback — same price, same lineage, one dropdown away.

Can I run Ember-1 locally? No. The launch materials announce no weights release; as of September 28, 2026 it’s hosted-only on Fireworks (plus OpenRouter and Vercel AI Gateway resale). The base model Kimi K3 has open weights under Modified MIT, but at 2.8T parameters it’s datacenter-scale either way.

Does Ember-1 work with Cursor Tab autocomplete? No — and neither does any BYOK model. Cursor’s documentation is explicit that custom API keys cover Chat and Agent only; Tab, Auto, Cloud/Background Agents, Automations, and the Cursor CLI always run on Cursor-routed models.

Should I use Ember-1 or Claude Fable 5.1 for agent sessions? Ember-1 for volume: a 50K-output-token session costs $0.75 versus $2.50 on Fable 5.1 before token reduction is counted. Fable 5.1 for the runs where a single point of SWE-bench Verified accuracy is worth 3x the money — typically large multi-file refactors you’ll only get one shot at reviewing.

Sources

Last updated September 28, 2026. Ember-1 is a research preview with a two-week serverless window — verify it is still served, and at these rates, before wiring it into anything permanent.

Was this article helpful?

Know which coding tool is worth paying for

Hands-on comparisons of AI coding assistants and what each one costs to run — including the local-model path. Sent only when something changes. Unsubscribe anytime.