As of: 2026-10-10 14:45 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/m/fireworks-ember-1/ # Ember-1 (Fireworks Research) Fireworks AI, Ember family. Status: current. Type: reasoning model. Released: 2026-09-23. Training cutoff: not published. ## How to call it | Provider | Model id | Endpoint | Docs | |---|---|---|---| | Fireworks AI | | | https://fireworks.ai/blog/ember-1 | | OpenRouter | `fireworks/ember-1` | | https://openrouter.ai/fireworks/ember-1 | ## Pricing - Input: $3 per 1M tokens - Cached input: $0.30 per 1M tokens - Output: $15 per 1M tokens - Unit as published: per 1M tokens (USD) - Source: https://openrouter.ai/fireworks/ember-1 - Last checked: 2026-10-10 ## Specs - Input: text - Output: text - Context window: 1,048,576 tokens - Max output: 943,718 tokens - Open weights: no - Licence: proprietary ## What it has not seen - No published training cutoff. Counted from its release (2026-09-23): at least 414 AI events since (75 major, 13 historic), as of 2026-10-10. - Everything since August 2026: https://postcutoff.com/since/2026-08/ ## What stands out - Shorter reasoning traces on Kimi K3: Post-trained from Kimi K3 to produce comparable answers with about 40% fewer tokens (Fireworks); SWE-bench Verified 92.2% vs 93.2% for K3, Terminal-Bench 2.1 82.0% vs 80.9%; one customer A/B test cut total tokens 39% and reasoning tokens 71.3% (vendor-reported, via Runtime Wire). Source: https://runtimewire.com/article/fireworks-ember-1-kimi-k3-reasoning-tokens ## Notes First model of Fireworks Research's Ember series, launched Sept 23, 2026 as a research preview and later listed as production-ready (Runtime Wire). Derivative of Moonshot's open-weight Kimi K3; no weights released. Fireworks' own blog post (fireworks.ai/blog/ember-1) was not read directly; prices and context from the OpenRouter API (checked 2026-10-10). Built for coding agents and multi-turn systems where reasoning tokens dominate cost. Fewer tokens lower per-task cost, but Fireworks' own caveat (per Runtime Wire) is that token savings alone do not prove lower latency across workloads. Model page: https://postcutoff.com/m/fireworks-ember-1/