Mercury 2.5
Artificial Analysis lists the release as Sept 8, 2026, measures ~661 tok/s and gives an Intelligence Index of 12, below average for its price tier. It lists $0.25/$0.75, while Inception's docs list $0.20/$0.75 before the discount. Inception's blog page showed Sept 29, 2026 when fetched, and OpenRouter also has a separate `inception/mercury-2.5-preview` listing, so the Sept 8 date may be the preview. Speed figures are vendor claims without disclosed batch or hardware details (RuntimeWire). Customer claims: Augment Code context compaction 150 s → 27 s; OpenCall ~170 ms median latency. 100M free API tokens for new accounts.
- Context window
- 260,000 tokens
- Max output
- 65,536 tokens
- Input
- text
- Output
- text
- License
- proprietary
- Pricing
- input: $0.2 · output: $0.75 (per 1M tokens (list price; 80% launch discount brings it to $0.04 / $0.15, end date not stated)) source
- Verified
- 2026-09-30
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Inception API (OpenAI-compatible) | mercury-2.5 | https://api.inceptionlabs.ai/v1/chat/completions | docs |
| OpenRouter | inception/mercury-2.5 | openrouter.ai/inception/mercury-2.5 | — |
| Baseten | — | www.baseten.co/ | — |
Notable capabilities (2)
- Diffusion-based text generation at very high speed: A diffusion LLM (dLLM) that refines many tokens in parallel rather than one at a time; Inception claims 1,107 tokens/s on NVIDIA GPUs, and Artificial Analysis measured ~661 tokens/s. source
- Fast reasoning with tool calling: Configurable reasoning effort, tool calling and structured outputs; Inception claims a 40% intelligence gain over Mercury 2 and places it near GPT-5.6 Luna (low), Gemini 3.5 Flash-Lite and Claude Haiku 4.5. source
Mercury 2.5 is aimed at latency-sensitive work (voice agents, autocomplete, agent sub-calls) where speed matters more than peak intelligence.
Sources: Inception blog · Artificial Analysis · RuntimeWire · Inception on X