{"schema":"postcutoff/model@1","as_of":"2026-10-10T14:45:00+02:00","url":"https://postcutoff.com/m/fireworks-ember-1/","md":"https://postcutoff.com/m/fireworks-ember-1/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"fireworks-ember-1","name":"Ember-1 (Fireworks Research)","org":"Fireworks AI","family":"Ember","released":"2026-09-23","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":943718,"knowledge_cutoff":null,"pricing":{"input":3,"cached_input":0.3,"output":15,"unit":"per 1M tokens (USD)","source":"https://openrouter.ai/fireworks/ember-1"},"price_line":"$3 in, $15 out per 1M tokens","access":[{"provider":"Fireworks AI","url":"https://fireworks.ai/blog/ember-1"},{"provider":"OpenRouter","model_id":"fireworks/ember-1","url":"https://openrouter.ai/fireworks/ember-1"}],"capabilities":[{"name":"Shorter reasoning traces on Kimi K3","detail":"Post-trained from Kimi K3 to produce comparable answers with about 40% fewer tokens (Fireworks); SWE-bench Verified 92.2% vs 93.2% for K3, Terminal-Bench 2.1 82.0% vs 80.9%; one customer A/B test cut total tokens 39% and reasoning tokens 71.3% (vendor-reported, via Runtime Wire).","first":false,"discovered":"launch","source":"https://runtimewire.com/article/fireworks-ember-1-kimi-k3-reasoning-tokens"}],"entry":null,"notes":"First model of Fireworks Research's Ember series, launched Sept 23, 2026 as a research preview and later listed as production-ready (Runtime Wire). Derivative of Moonshot's open-weight Kimi K3; no weights released. Fireworks' own blog post (fireworks.ai/blog/ember-1) was not read directly; prices and context from the OpenRouter API (checked 2026-10-10).","verified":"2026-10-10","body_md":"Built for coding agents and multi-turn systems where reasoning tokens dominate cost. Fewer tokens lower per-task cost, but Fireworks' own caveat\n(per Runtime Wire) is that token savings alone do not prove lower latency across workloads.","page_url":"https://postcutoff.com/m/fireworks-ember-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":414,"you_url":"https://postcutoff.com/you/fireworks-ember-1/","briefings":null}