News
Fireworks Post-Trained Kimi K3 to Think 40% Less (Ember-1, 2026)
The fastest way to cut AI agent bills in 2026 is making the model shut up. Fireworks AI post-trained Moonshot's 2.8T Kimi K3 into Ember-1, producing identical answer quality with 40 percent fewer reasoning tokens at the same $3/$15 sticker price. In multi-turn coding loops where verbose thoughts compound into subsequent input context, that reduction protects SaaS margins from terminal decline. Here is the founder math behind the release, the double-billing mechanism, and where startup directories like SaaSCity fit into the launch.

Contents (8)
- The spec sheet: Kimi K3 under the hood
- Inside the post-training process
- The double-billing trap in multi-turn agent workflows
- Client routing versus model weights
- The Hacker News reaction and the post-training playbook
- How to benchmark Ember-1 on your own workload
- Production caveats: what the research preview label means
- The bottom line on token economics
Quick answer: On September 23, 2026, Fireworks AI released Ember-1, a specialized model built on Moonshot AI's 2.78-trillion parameter Kimi K3. By post-training the base weights to produce compact reasoning traces, Ember-1 matches K3 quality while consuming roughly 40 percent fewer tokens. With list pricing held flat at $3.00 input, $0.30 cached input, and $15.00 output per million tokens, the financial win comes from curbing token inflation in multi-turn agent workflows.

Start here: if you are building an AI software startup or developer tool, list your product on SaaSCity. You are reading this analysis on our directory's publication, so consider this an honest early disclosure. SaaSCity is a gamified startup directory with a live city map and human editorial review. A free submission gives you a permanent listing page and a building on the interactive map. When you add the SaaSCity badge to your website, you unlock a dofollow backlink and secure a Monday launch slot. Founders looking to skip the badge can pick Quick Pass for $19.99 to go live within 24 hours, while the Premium tier at $99.99 adds a written launch post with three dofollow links. Our domain sits between DR 47 and DR 56 on recent Ahrefs crawls.
The fastest way to cut your AI bill in 2026 is making the model shut up.
For twelve months, software teams treated reasoning traces as free intellectual equity. Whenever an evaluation stalled, engineering leads dialed up reasoning effort, letting chains of thought balloon into tens of thousands of tokens. But in multi-turn coding agents, a verbose reasoning trace is billed as output at $15.00 per million tokens on turn one, and re-sent as input on turns two through ten. Verbosity became a compounding tax on gross margins.
On September 23, 2026, Fireworks AI pushed an answer to that balance sheet problem. They announced Introducing Ember-1, the first custom model series from Fireworks Research. Ember-1 keeps the standard sticker price but slashes the volume of tokens required to reach the right answer by roughly 40 percent. Four days later, the release reached the front page of Hacker News. The Hacker News discussion accumulated 508 points and 225 comments, focusing on founder arithmetic: saving up to $68 per software patch.
The spec sheet: Kimi K3 under the hood
Moonshot AI released Kimi K3 on July 16, 2026 as an open-weight mixture-of-experts model spanning 2.78 trillion parameters. Its weights were published on July 27 under a modified-MIT license, as documented in the CodersEra guide to Kimi K3. The architecture routes tokens through 16 active experts out of 896 total sub-networks, pairing world knowledge with a context window of 1,040,000 tokens.

Kimi K3 earned adoption among software developers for coding and logic accuracy. Yet thinking mode is always on, and reasoning effort runs at maximum by default. While this produces thorough answers, it generates sprawling internal monologues. Basic coding requests trigger thousands of reasoning tokens before the model writes its first line of code.
According to the official Ember-1 model card, created on September 22, 2026 with model path accounts/fireworks/models/ember-1, Fireworks took Moonshot's base weights and applied targeted post-training. The model card lists Ember-1 as a specialized base model with 2.78 trillion parameters, function-calling support, a 1.04-million token context window, and availability via Fireworks Serverless.
The official performance table published by Fireworks Research outlines what this post-training produced:
| Metric | Baseline Kimi K3 Configuration | Fireworks Ember-1 | Measured Change |
|---|---|---|---|
| Benchmark Accuracy Score | 0.753 | 0.753 | Parity maintained |
| Average Execution Steps | 21.4 | 21.4 | Parity |
| Total Output Tokens Generated | 49.0K | 29.9K | 39.0% reduction |
| Internal Reasoning Tokens | Baseline (100%) | 28.7% of baseline | 71.3% reduction |
| SWE-bench Cost per Task | Full compute cost | Baseline minus up to $68 | Up to $68 saved per task |
The standout figure is the 71.3 percent drop in hidden reasoning tokens. Across programming workflows, Ember-1 skips circular self-questioning and redundant verification passes. Total generated tokens fell by 39 percent while preserving an identical accuracy score of 0.753.
Benchmark figures collected by the HyperAI release summary indicate that total token reductions consistently ranged between 35 and 50 percent across coding suites. On SWE-bench Verified, Ember-1 matched Kimi K3's maximum-reasoning success rate while strictly dominating the base model's lower-effort settings. When developers manually lowered Kimi K3's reasoning effort to save money, accuracy fell off. Ember-1 preserved high accuracy while operating at lower token volumes.
On the Specialized Intelligence Index, Ember-1 registered comparable gains. On Doximity's Bedside Bench, a rigorous medical logic evaluation, Ember-1 set a new Pareto frontier for cost versus task completion, beating proprietary models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on completed cost per task.
Inside the post-training process
Trimming reasoning tokens without degrading problem-solving ability is difficult. Fine-tuning models on short answers often causes catastrophic forgetting: models stop thinking through hard edge cases, producing hallucinations instead of verified solutions.
Fireworks conducted over 50 discrete training experiments and ran more than 200 automated evaluation cycles to calibrate their reward formulas. The curriculum ran on Fireworks Serverless Training infrastructure, targeting five core domains: software engineering, advanced mathematics, multi-turn instruction following, tool use, and extended agentic interactions.
The training algorithms targeted unproductive reasoning loops. If the model caught a bug on reflection step one, it received reward for executing immediately instead of generating three more paragraphs of confirmatory thoughts.
Fireworks tested the model internally by dogfooding it across its engineering team. Developers wrote features, patched services, and ran test suites with daily coding assistants redirected to Ember-1. The team reported zero difference in code quality or completion success while backend token meters dropped significantly. Internal leadership summarized the rollout simply: "no news is good news."
Subsequent production trials with enterprise beta customers showed identical outcomes. Customer coding workloads measured an average 35 percent drop in total token consumption with flat or slightly improved task completion rates. One enterprise customer confirmed plans to decommission their existing frontier base models in favor of Ember-1 endpoints.
The double-billing trap in multi-turn agent workflows
To understand why a 40 percent volume reduction matters more than a discount on API prices, examine how modern developer tools consume tokens.

The OpenRouter Ember-1 listing, released on September 24, 2026, lists input at $3.00 per million, cached input at $0.30, and output at $15.00 across a 1.0M context window. Observed throughput runs 49 to 69 tokens per second (52 tps median) with a 0.65-second time-to-first-token, a 33.7-second P50 end-to-end latency, and 98.95 percent availability.
Fireworks did not discount the nominal rates. Ember-1 costs exactly what Kimi K3 costs across commercial hosts. According to independent analysis from the Mercatus AI pricing report, which tracks Kimi K3 across 11 hosting platforms, K3's $15.00 output rate makes it the most expensive open-weight model on the market, roughly 3.4 times higher than the output fees for GLM-5.2.
The savings come from eliminating what tech analyst Jim Clyde Monge described in his Zeniteq breakdown as the double-billing mechanism of autonomous agents.
Consider an agent tasked with refactoring an authentication module across a repository over ten turns. In an agent loop, you pay for reasoning tokens twice. First, you pay for them at the steep generation rate of $15.00 per million. Second, your client library packs those same reasoning tokens into the message payload on every subsequent turn. Even with prompt caching enabled, those tokens incur a persistent carrying fee. If prompt caching misses, you pay the full $3.00 per million input rate on every turn.
Here is the exact math from Zeniteq's worked example:
Suppose Ember-1 cuts 10,000 reasoning tokens per turn compared to the base model. Across a 10-turn task, the immediate output savings are straightforward: 10 turns multiplied by 10,000 tokens equals 100,000 output tokens saved. At $15.00 per million, that shaves $1.50 off your output bill immediately.
Now examine the downstream input savings across those ten turns. On turn one, zero past tokens exist. On turn two, you avoid sending 10,000 past tokens. On turn three, you avoid 20,000 tokens, continuing through turn ten where you avoid 90,000 accumulated tokens. The cumulative avoided input context equals 450,000 tokens (10,000 + 20,000 + ... + 90,000).
Depending on your cache architecture, the input savings break down as follows:
- Uncached input ($3.00 / 1M): 450,000 multiplied by $0.000003 equals $1.35 saved.
- Cached input ($0.30 / 1M): 450,000 multiplied by $0.0000003 equals $0.135 saved.
Adding output and input savings together, a single 10-turn software engineering task saves between $1.635 and $2.85.
| Billing Component | Base Kimi K3 (Verbose) | Fireworks Ember-1 (Trimmed) | Net Task Savings |
|---|---|---|---|
| Turn 1-10 Output (at $15/1M) | 150,000 output tokens ($2.25) | 50,000 output tokens ($0.75) | $1.50 |
| Cumulative Context Input (Uncached at $3/1M) | 675,000 input tokens ($2.025) | 225,000 input tokens ($0.675) | $1.35 |
| Cumulative Context Input (Cached at $0.30/1M) | 675,000 input tokens ($0.202) | 225,000 input tokens ($0.067) | $0.135 |
| Total Cost per 10-Turn Task (Uncached) | $4.275 | $1.425 | $2.850 (66.7% cut) |
| Total Cost per 10-Turn Task (Cached) | $2.452 | $0.817 | $1.635 (66.7% cut) |
When your application processes 5,000 coding tickets a day, that difference amounts to $8,175 to $14,250 in cash saved every day. The gross margin on your software feature flips from a deficit into positive territory.
Client routing versus model weights
Trimming token spend is the primary engineering objective for developer platforms in 2026. Teams attack the issue at different layers of the software stack.
Earlier this month, we covered how Spotify addressed cost pressures in our report on how Spotify cut Claude Code token costs 90% with a two-model router. Spotify's engineering team built a client-side routing plugin called shunt. Instead of modifying the foundation model, they used hooks to intercept large file reads, delegating bulk scanning to a lightweight worker model while reserving Claude for synthesis.
Fireworks took an alternate path. Rather than introducing orchestration complexity, custom middleware, and secondary models, Ember-1 attacks token bloat inside the model weights. The developer does not need to configure hooks or write routing heuristics. You point your existing OpenAI-compatible client at accounts/fireworks/models/ember-1, and the model stops generating redundant tokens on its own.
Both approaches have merits, and sophisticated platforms are beginning to combine them. Teams running complex agent systems must balance model weights against their underlying execution environments, as analyzed in our review of agent harness costs across AWS, Strands, and DigitalOcean. If your harness environment costs $0.05 per minute to keep active while waiting for slow inference streams, Ember-1's shorter generations lower both your token invoice and your container runtime bill.
At the same time, hardware and open-weight dynamics continue to shift globally. In our breakdown of Xiaomi MiMo-V2.6 open weights model economics, we documented how Chinese labs are providing enterprise-scale reasoning at commodity pricing. Ember-1 demonstrates how Western hosting providers can take those open weights, apply specialized post-training, and build proprietary competitive advantages.
The Hacker News reaction and the post-training playbook
When the Ember-1 announcement hit Hacker News on September 27, the developer response highlighted a turning point in how technical teams view machine learning infrastructure.

The top-voted comment on the thread, posted by user GodelNumbering, captured the prevailing sentiment:
"This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!"
Other engineers shared parallel experiences. Developers are recognizing that general-purpose foundation models are inherently inefficient for specific domain tasks. A frontier model trained to write poetry, summarize legal filings, and write Python code wastes capacity on general reasoning steps when tasked with a narrow coding objective.
This community reaction explains why teams are feeling overwhelmed by the endless cycle of new frontier model releases. We explored this trend in our guide on combating model fatigue when choosing AI models for SaaS. For two years, the standard playbook was reactive: every time a major lab dropped a new checkpoint, developers rushed to swap API keys, rewrite system prompts, and hope their unit economics held up.
Ember-1 represents a proactive alternative: post-training open weights for behavioral efficiency. Rather than switching between rival foundation giants, as we explored when analyzing Kimi K3 vs Qwen3.8-Max in China's trillion-parameter model war, companies can take an established open-weight model and train it to adhere to strict corporate brevity standards.
Recognizing this market shift, Fireworks is not keeping this post-training capability locked away as an internal secret. Along with the Ember-1 release, Fireworks rolled out enterprise support for its Serverless Training infrastructure. Any engineering organization can upload proprietary interaction logs, synthetic agent traces, and internal code reviews to produce custom, token-efficient model variants tuned to their company's operational needs.
How to benchmark Ember-1 on your own workload
Before switching production endpoints to Ember-1, set up an objective testing harness rather than relying solely on published vendor figures.
Start by building a frozen evaluation bank. Extract 100 to 200 real user interactions from your platform logs. Ensure the sample includes multi-file bug fixes, greenfield feature builds, complex refactors, and simple single-turn syntax queries. Strip all user identifying data and store these traces as your permanent benchmark suite.
Next, measure pass rates alongside token volume. Run baseline Kimi K3 and Ember-1 across the frozen dataset under identical temperature and top-p parameters. Track four numbers for each run: task success against your test suite, total output tokens, hidden reasoning tokens, and total wall-clock duration. Calculate your effective cost per resolved issue: total dollars spent divided by successful task completions. A model that cuts token consumption by 40 percent but suffers a 10 percent drop in test pass rates forces costly retries that erase paper savings.
Audit failure modes carefully. Examine the specific prompts where Ember-1 failed while base Kimi K3 succeeded. When reasoning models are trained to think less, their most frequent regression point is edge-case exploration. Watch for omitted error handling, race conditions, or incomplete code stubs. If regressions concentrate in complex domains your product touches daily, consider keeping base Kimi K3 for high-difficulty tickets while routing standard tickets to Ember-1.
Finally, monitor real-world latency dynamics. OpenRouter metrics for Ember-1 show a median time-to-first-token of 0.65 seconds, but a median end-to-end duration of 33.7 seconds. Generating code through 16 active experts across a 2.78-trillion parameter MoE model requires significant memory bandwidth. If your user interface requires rapid conversational replies, a 30-second completion time might challenge user patience. Verify that generation pacing aligns with your user expectations.
Production caveats: what the research preview label means
While the economic results behind Ember-1 are notable, engineering leaders should evaluate several practical caveats before wiring it into mission-critical architecture.
All published performance numbers come from Fireworks' internal evaluations. As of September 28, 2026, no independent third-party evaluation group or academic benchmark collective has published external figures. While customer testimonials and the HyperAI summary corroborate the general 35 to 40 percent reduction range, independent verification remains pending.
Ember-1 does not ship with open weights. Moonshot released Kimi K3's weights under a modified-MIT license, giving organizations the legal right to host, fine-tune, and run the model on private GPU clusters. Fireworks has not released weights for Ember-1. There is no Hugging Face model repository, no Ollama manifest, and no downloadable safetensors checkpoint. Adopting Ember-1 binds your software to Fireworks' cloud infrastructure.
Availability is currently classified as a Research Preview. Technical reports from Zeniteq indicate that serverless access may be subject to a time-limited evaluation window of roughly two weeks before transitioning to paid on-demand enterprise tiers. If your architecture requires long-term service level agreements, confirm availability terms with Fireworks before committing production agent fleets to the model path.
Actual financial savings will depend on prompt caching efficiency. If your agent harness already maintains a 90 percent cache hit rate across conversation histories, input savings from trimmed reasoning tokens will be modest ($0.135 per 10-turn session rather than $1.35). As shown in the billing table, the primary driver of savings is output generation. If your tasks are input-heavy reads with very short outputs, the financial impact will be smaller than SWE-bench results suggest.
The bottom line on token economics
The release of Ember-1 marks a welcome transition in machine learning: industry attention is shifting away from raw parameter scale toward inference efficiency.
For three years, foundation labs operated under the assumption that longer reasoning traces were always preferable. Model providers billed users by the token, creating an obvious conflict of interest where more verbose responses generated higher cloud revenue.
Fireworks showed that post-training open weights to eliminate reasoning bloat delivers tangible margin relief to software developers. An application that cuts 40 percent of its token volume without sacrificing problem-solving quality can double its gross margins without waiting for GPU hardware price drops.
If you run software engineering agents or complex reasoning workflows, test Ember-1 on your frozen prompt suite today. If accuracy holds up on your specific domain tasks, you can secure meaningful margin expansion immediately. For teams with proprietary datasets, the broader lesson is clear: future AI economics favor post-training open models to solve company-specific tasks in as few tokens as possible.
When you have built that product and are ready to put it in front of founders and technical buyers, claim your space on SaaSCity and let the ecosystem see what you have shipped.
Get your SaaS in front of founders
List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.


