Ember-1: Kimi K3 Quality at Half the Tokens
Fireworks Research released Ember-1, a specialized model built on Kimi K3 that delivers the same quality answers while using 40% fewer tokens. It is the first in a planned series of specialized models from Fireworks, and it is available now as a research preview on their serverless platform.
The core problem is one that anyone running reasoning models at scale knows well: thinking models think too much. Kimi K3 sometimes spends more than 90% of its generated tokens on internal reasoning rather than the answer. In multi-turn agentic workloads, this gets worse because every turn replays all prior reasoning back to the model. Context grows quadratically. Cost grows with it.
Fireworks found that simply turning down K3's reasoning effort did not work. Lower settings gave up too much quality. So they trained the model to reason more efficiently instead. The team ran more than 50 training experiments and over 200 evaluations, developing new training algorithms to shorten reasoning without losing accuracy. They trained across mathematics, coding, instruction following, conversation, search, tool use, and software engineering.
The results are strong. On SWE-bench Verified, Ember-1 scored 92.2% compared to K3 max's 93.2%, but at 15.5% lower cost. On Terminal Bench 2.1, it actually scored higher (82.0% vs 80.9%) while cutting cost by 51.9%. On DeepSWE 1.1, it jumped from 66.4% to 75.2% at 23.7% lower cost. The token savings held up across seven benchmarks and two customer A/B tests.
Fireworks also evaluated Ember-1 on the Specialized Intelligence Index, benchmarking it against GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost per task using Doximity's Bedside Bench, a physician-validated benchmark spanning 500 clinical cases. Ember-1 set a new Pareto frontier, meaning no other model matched its combination of quality and cost.
The internal validation is the part I find most convincing. Fireworks replaced Kimi K3 with Ember-1 on their own developers' coding workloads without telling them. Nobody noticed. The developers carried on their work while consuming substantially fewer tokens. For a model whose pitch is "same answers, fewer tokens," an invisible rollout is the strongest possible signal.
Ember-1 is available as a two-week research preview on Fireworks Serverless. If it gets adoption, it becomes permanent. For anyone running K3 in production coding or agentic workflows, this is worth testing.
Source: Fireworks AI Blog [1]
Hacker News discussion (400 points, 199 comments) [2]
Ember-1 Model Page [3]