Fireworks Taught Kimi K3 to Stop Rambling — and the AI Cost Crisis Just Got Its First Real Fix


The bill nobody talks about is the one for thinking

Every AI company advertises the tokens you see: the answers, the code, the summaries. But for reasoning models — the ones that actually do hard work — the visible output is the tip of the invoice. The real cost is the thinking: long internal reasoning traces that can consume more than 90% of generated tokens before the model ever writes the first line of its answer.

That cost compounds viciously in agentic workloads. In a multi-turn coding agent, every turn replays all prior reasoning back into context. Context grows roughly quadratically with turn count, which means a reasoning trace that cost ten cents on turn one gets re-read and re-billed on turn two, turn three, and every turn after. The industry has known about this for a while. What it hasn’t had is a fix — until this week.

On September 23, Fireworks AI announced Ember-1: a specialized model from its research team that delivers Kimi K3’s quality with roughly 40% fewer tokens. It is, as far as the market has produced, the first serious attempt to attack reasoning-token bloat at the model level rather than with pricing games or API knobs.

Why the obvious fix didn’t work

The first thing teams tried, naturally, was the knob. Kimi K3 and other reasoning models expose reasoning-effort settings; turn the effort down, fewer tokens, lower bill. Fireworks says its customers tried exactly this and it failed in practice: cutting effort cut quality too much to be worth it for teams that wanted K3-level coding performance.

That outcome is worth sitting with, because it kills a tempting assumption. Reasoning effort isn’t a volume dial where less thinking means proportionally worse answers — it turns out to be a much worse trade than expected. The reasoning a model does isn’t all equally valuable. Some of it is load-bearing: the model catching its own bad assumption mid-answer, double-checking a tricky branch, exploring a genuinely ambiguous problem. Some of it is pure waste: repetitive loops, redundant self-checks, the model re-deriving something it already established.

The knob can’t tell the difference. It just does less of everything. What Ember-1 represents is the recognition that efficiency has to be learned, not dialed.

What Fireworks actually did

Ember-1 is not a new foundation model. Fireworks took Moonshot AI’s open-weight Kimi K3 and ran it through additional reinforcement learning aimed at a single behavior: reach the same answer with a shorter reasoning trace. Keep the thinking that matters; strip the loops that burn tokens.

Getting there took more than a tweak. Fireworks Research ran more than 50 training experiments and 200-plus evaluations across math, coding, tool use, and software engineering, developing new training algorithms along the way to shorten reasoning without losing accuracy. The whole program ran on Fireworks’ own serverless training platform — which is itself a quiet product pitch: the company needed no GPU provisioning or cluster management to go from research idea to launched model.

The reported numbers are strong, with the usual vendor-reporting caveats. On coding benchmarks, Ember-1 holds within about a point of Kimi K3 Max on SWE-bench Verified (92.2% vs. 93.2%) and SWE-Interact (20.0% vs. 21.3%), while actually beating K3 Max on Terminal Bench 2.1 (82.0% vs. 80.9%) and DeepSWE 1.1 (75.2% vs. 66.4%). Token savings vary by workload: 15.5% fewer tokens on SWE-bench Verified, 51.9% on Terminal Bench 2.1, and roughly 40% on average across the board.

More interesting than the benchmarks: Fireworks reports live A/B tests with two production coding customers showing approximately 35% fewer tokens per task at comparable quality, and claims its own developers ran Ember-1 internally for a stretch without noticing it had replaced K3. An independent check from one coverage outlet found the vendor’s K3 baseline numbers unusually tight with third-party leaderboard scores — which doesn’t validate Ember-1’s claims directly, but suggests the yardstick isn’t rigged.

Pricing stays at the public Kimi K3 rate — $3 per million input tokens, $15 per million output. Since you’re simply billed for fewer tokens per task, the bill shrinks without any price-sheet maneuvering.

Why this matters beyond one model

Three things about this release signal something bigger than a single model launch.

First, efficiency is becoming a product category, not a config option. Fireworks explicitly says Ember-1 kicks off an ongoing series of specialized models from Fireworks Research, shaped by developer demand. That’s a meaningful strategic bet: that the market will pay for models differentiated by cost-per-task rather than raw capability. It also sets up a new competitive axis — not who has the smartest model, but who has the smartest model per dollar of inference.

Second, it reframes where the AI cost crisis gets solved. The dominant conversation about AI economics has been about GPUs, data centers, and power. But the marginal cost that decides whether agentic products are viable businesses is inference cost per task — and the biggest lever on that cost is how much a model talks to itself. If training can systematically compress reasoning traces without losing quality, the economics of agent-heavy products change materially. Coding agents, research agents, support agents — all of them spend their budget on internal monologue.

Third, it arrives at exactly the right moment. In the same week, OpenAI shelved its next flagship model over safety concerns and paused frontier training. When frontier capability timelines get murky, efficiency improvements are the gains the industry can still count on. A 40% cost cut on tasks you run today is worth more than a hypothetical model that may or may not ship next quarter. The teams that built their roadmaps around models arriving on schedule just learned a lesson; the teams that built around cost-per-task just got a gift.

The caveats, stated plainly

The honest fine print: Ember-1 is currently available only as a Research Preview through Fireworks’ serverless API — and through Vercel’s AI Gateway as fireworks/ember-1. No weights, training code, or algorithm details have been released, so self-hosting isn’t an option. That means the efficiency techniques stay proprietary; you can rent the result, not replicate it.

The benchmark numbers are vendor-reported on the vendor’s own harness for the headline claims, and “40% fewer tokens” is a workload-dependent average, not a guarantee. Any team should run its own evaluation on its own tasks before renegotiating its cost models — particularly because token savings varied from 15% to over 50% across benchmarks.

And there’s a subtler question worth asking: when a model learns to hide its reasoning, does auditability suffer? This is the same week the industry is reckoning with agents that misrepresent what they did. Shorter reasoning traces are cheaper — but they also give humans less to inspect when something goes wrong. For deployments where you need to see the work, the trace compression that makes Ember-1 attractive may need to be weighed against observability requirements. Efficiency and transparency pull in opposite directions here, and nobody has resolved that tension yet.

The bottom line

Ember-1 won’t make headlines the way a frontier launch does. It’s a cost model, not a capability leap. But cost models are what turn AI from a demo into a business. The reasoning-token tax has been the industry’s open secret — the reason agents that work brilliantly in demos get expensive in production. This is the first credible, trained-into-the-model fix for it, and it establishes a new idea of what progress looks like: not just models that think better, but models that think less.

If you’re running coding agents or multi-turn workloads in production, the practical move is simple: evaluate Ember-1 on your actual tasks in the preview window, measure tokens-per-task on your own traffic, and let your bill decide. The knob didn’t work. The training might.