# AI model registry (Post-Cutoff)
Generated 2026-09-29. 270 models. Verify prices and ids against the linked docs before production use.

## 1X Technologies

### 1X Redwood AI
**1X Redwood AI** (1X Technologies; current; robotics; released 2025-06-10) | Onboard 1X NEO (consumer humanoid, preorder): https://www.1x.tech/order — Ships as NEO's foundational autonomy; tasks it cannot do are handled by remote human teleoperation, which drew privacy criticism (https://startupfortune.com/1xs-20000-neo-robot-lets-a-company-employee-watch-inside-your-home/). NEO: $20,000 Early Access ownership or $499/month subscription, $200 refundable deposit, "US deliveries start 2026" (order page checked 2026-09-29). 1X opened its Hayward, CA NEO factory on 2026-04-30 (10,000 units targeted in year one); as of mid-July 2026 no verified customer home delivery had been reported and we found none by 2026-09-29. See also 1x-world-model (video world-model policy, Jan 2026).
  - Onboard 1X NEO (consumer humanoid, preorder) — https://www.1x.tech/order (docs: https://www.1x.tech/discover/redwood-ai)
Notable capabilities:
  - Small onboard VLA for a home humanoid: 160M-parameter vision-language transformer (language embeddings + ViT tokens + proprioception) with a diffusion-policy action decoder, running fully on NEO's embedded GPU at ~5 Hz, so it works without internet. (https://www.1x.tech/discover/redwood-ai)
  - Mobile bimanual whole-body manipulation: Combines locomotion with manipulation (bending, leaning, bracing) for retrieving objects, opening doors and navigating the home; learns from both successful and failed episodes. (https://www.1x.tech/discover/redwood-ai)
  - Voice control via offboard LLM: An offboard speech-to-speech LLM handles conversation and hands tasks to Redwood. (https://www.1x.tech/discover/redwood-ai)

### 1X World Model (1XWM)
**1X World Model (1XWM)** (1X Technologies; preview; world-model; released 2026-01-12) | Not available (internal; runs NEO policies): https://www.1x.tech/discover/world-model-self-learning — Two stages: 1XWM as a policy evaluator (2025-06-16) and as a NEO policy (2026-01-12). No API or weights. TechCrunch coverage: https://techcrunch.com/2026/01/13/neo-humanoid-maker-1x-releases-world-model-to-help-bots-learn-what-they-see/
  - Not available (internal; runs NEO policies) — https://www.1x.tech/discover/world-model-self-learning (docs: https://www.1x.tech/1x-world-model.pdf)
Notable capabilities:
  - Video world model used as the robot policy: Given a text prompt, a 14B generative video model fine-tuned on NEO imagines ~5 s of future video; an inverse-dynamics model converts it into actions executed on NEO (≈11 s per rollout on multi-GPU inference). (https://www.1x.tech/discover/world-model-self-learning)
  - Learns from human egocentric video: Trained with ~900 h of egocentric human video plus ~70 h of NEO data (and 400 h of unfiltered robot data for the IDM); generalizes to some objects and motions absent from NEO task data. Grasping ~80% success; pouring 0%; best-of-8 generations raised 'pull tissue' from 30% to 45%. (https://www.1x.tech/discover/world-model-self-learning)
  - World model for policy evaluation: The June 2025 version was an action-conditioned simulator used to rank policies without physical tests (1X: 70% world-model accuracy picks the better policy ~90% of the time). (https://www.1x.tech/discover/redwood-ai-world-model)


## ACE Studio & StepFun

### ACE-Step 1.5 (incl. 1.5 XL)
**ACE-Step 1.5 (incl. 1.5 XL)** (ACE Studio & StepFun; current; music; released 2026-01-28; open weights) | Hugging Face: https://huggingface.co/ACE-Step/Ace-Step1.5; Hugging Face (XL 4B DiT): https://huggingface.co/ACE-Step/acestep-v15-xl-sft; GitHub: https://github.com/ace-step/ACE-Step-1.5; Web app: https://acemusic.ai — Checkpoints: acestep-v15-base / -sft / -turbo (plus turbo-shift variants) and, from 2026-04-02, XL (4B DiT) xl-base / xl-sft / xl-turbo; diffusers versions added Apr-Jun 2026. Release date 2026-01-28 is from secondary sources (HF repos created 2026-01-23, arXiv 2602.00744 submitted 2026-01-31). Authors claim quality beyond most commercial models (SongEval above Suno v5 per secondary coverage; not independently verified). Supports Mac, AMD, Intel and CUDA.
  - Hugging Face — https://huggingface.co/ACE-Step/Ace-Step1.5
  - Hugging Face (XL 4B DiT) — https://huggingface.co/ACE-Step/acestep-v15-xl-sft
  - GitHub — https://github.com/ace-step/ACE-Step-1.5
  - Web app — https://acemusic.ai
Notable capabilities:
  - Full songs in seconds on consumer hardware: 10 s to 10 min of music; under 2 s per song on an A100 and under 10 s on an RTX 3090; standard models run in <4 GB VRAM with offload (XL: >=12 GB, 20 GB recommended). (https://github.com/ace-step/ACE-Step-1.5)
  - LM planner + DiT synthesizer: A language model (0.6B/1.7B/4B '5Hz LM') turns prompts into a song blueprint that a Diffusion Transformer renders; aligned with 'intrinsic' RL without external reward models. (https://arxiv.org/abs/2602.00744)
  - Editing and personalization toolkit: Cover generation, repaint/editing, vocal-to-BGM, track separation, multi-track generation, BPM/key extraction and LoRA fine-tuning from ~8 songs (about 1 h on a 12 GB RTX 3090); lyrics in 50+ languages. (https://github.com/ace-step/ACE-Step-1.5)


## AgiBot

### AgiBot GO-2 (Genie Operator-2)
**AgiBot GO-2 (Genie Operator-2)** (AgiBot; current; robotics; released 2026-04-09) | AgiBot robots / Genie Studio (via AgiBot sales): https://www.agibot.com/article/231/detail/56.html — No open weights, API or pricing found (GO-1 was open, non-commercial). Core work accepted to CVPR 2026 and ACL 2026 per AgiBot. Trained on 'tens of thousands of hours' of interaction data.
  - AgiBot robots / Genie Studio (via AgiBot sales) — https://www.agibot.com/article/231/detail/56.html
Notable capabilities:
  - Action chain-of-thought: Reasons in action space: generates a macro-plan of high-level action intents, then executes step by step, with teacher forcing so execution adheres to the reasoning. (https://www.agibot.com/article/231/detail/56.html)
  - Asynchronous dual-system: Low-frequency semantic planner ('commander') plus high-frequency action follower ('executor') in one architecture. (https://www.therobotreport.com/agibot-releases-go-2-foundation-model-embodied-ai/)
  - Benchmark results: LIBERO 98.5% average (ranked 1st), LIBERO-Plus 86.6% zero-shot, VLABench 47.4, 82.9% real-world success from simulation-only training (company-reported). (https://www.agibot.com/article/231/detail/56.html)

### AgiBot GO-1 (Genie Operator-1)
**AgiBot GO-1 (Genie Operator-1)** (AgiBot; legacy; robotics; released 2025-03-10; open weights) | Hugging Face: `agibot-world/GO-1`; Hugging Face (lighter variant): `agibot-world/GO-1-Air` | GitHub: https://github.com/OpenDriveLab/Agibot-World — Paper arXiv 2503.06669 (2025-03-09); announced ~2025-03-10 (day not re-verified). Weights on HF from Sept 2025, non-commercial license. Successor: agibot-go-2 (Apr 2026).
  - Hugging Face: `agibot-world/GO-1` — https://huggingface.co/agibot-world/GO-1
  - Hugging Face (lighter variant): `agibot-world/GO-1-Air` — https://huggingface.co/agibot-world/GO-1-Air
  - GitHub — https://github.com/OpenDriveLab/Agibot-World
Notable capabilities:
  - Latent-action VLA trained on AgiBot World: 3B model on an InternVL2.5-2B backbone using latent action representations, pretrained on AgiBot World (1M+ trajectories, 217 tasks, 5 deployment scenarios); ~30% average gain over policies trained on Open X-Embodiment, 60%+ success on complex tasks, +32% vs RDT. (https://arxiv.org/abs/2503.06669)


## Agility Robotics

### Agility Digit whole-body control foundation model ("motor cortex")
**Agility Digit whole-body control foundation model ("motor cortex")** (Agility Robotics; current; robotics; released 2025-08-28) | Onboard Agility Digit (commercial humanoid, via Agility): https://www.agilityrobotics.com/content/agility-and-ai — Agility has not published a large VLA of its own; this is its disclosed foundation-model layer. Digit is in paid deployments (e.g. GXO); Agility opened a Fremont "Physical AI" facility in July 2026 (https://www.nasdaq.com/press-release/agility-opens-new-fremont-facility-accelerate-physical-ai-development-2026-07-16). Not developer-accessible.
  - Onboard Agility Digit (commercial humanoid, via Agility) — https://www.agilityrobotics.com/content/agility-and-ai (docs: https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model)
Notable capabilities:
  - Tiny sim-trained whole-body controller: An LSTM with fewer than 1M parameters, trained with RL in NVIDIA Isaac Sim for decades of simulated time in 3-4 days, transferring zero-shot to Digit for balance, walking, arm placement and carrying heavy objects while staying stable. (https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model)
  - Layered stack with LLM on top: Higher layers (open-vocabulary detectors, state-machine planners, an LLM such as a Gemini research preview) send targets to the motor cortex; dexterous skills are learned on top of it. (https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model)


## Ai2 (Allen Institute for AI)

### MolmoAct 2 / MolmoAct 2-Think
**MolmoAct 2 / MolmoAct 2-Think** (Ai2 (Allen Institute for AI); current; robotics; released 2026-05-05; open weights) | Hugging Face: `allenai/MolmoAct2`; Hugging Face LeRobot: `allenai/MolmoAct2-LIBERO-LeRobot` | GitHub: https://github.com/allenai/molmoact2 — Checkpoints: MolmoAct2 (post-trained multi-embodiment foundation, ~5.4B params per HF safetensors), -Think, -Pretrain, fine-tuned -DROID, -BimanualYAM, -SO100_101, -LIBERO, -Think-LIBERO, FAST-Tokenizer. Main supported robots: SO-100/101, bimanual YAM, Franka (DROID); others need fine-tuning. Paper arXiv 2605.02881.
  - Hugging Face: `allenai/MolmoAct2` — https://huggingface.co/collections/allenai/molmoact2-models
  - GitHub — https://github.com/allenai/molmoact2
  - Hugging Face LeRobot: `allenai/MolmoAct2-LIBERO-LeRobot` — https://huggingface.co/allenai/MolmoAct2-LIBERO-LeRobot
Notable capabilities:
  - Open action reasoning model: Molmo2-ER embodied-reasoning VLM connected to a flow-matching action expert via per-layer KV conditioning; the Think variant adds adaptive depth reasoning (interpretable depth map before acting). (https://allenai.org/blog/molmoact2)
  - Strong out-of-the-box real-world success: 87.1% average success over 15 real Franka tasks vs 45.2% for π0.5 and 48.4% for MolmoBot (Ai2's evaluation); LIBERO 97.2% (98.1% Think). (https://allenai.org/blog/molmoact2)
  - Fast inference: ~180 ms per action call (790 ms with adaptive depth reasoning) vs ~6,700 ms for the original MolmoAct (up to 37x faster). (https://allenai.org/blog/molmoact2)
  - Largest open bimanual dataset: Released with MolmoAct2-BimanualYAM, 720+ hours of bimanual tabletop demonstrations, which Ai2 calls the largest open bimanual robotics dataset, plus an open FAST action tokenizer. (https://allenai.org/blog/molmoact2)


## Alibaba (Qwen / Tongyi Lab)

### Qwen-Audio-3.0-TTS (Flash / Plus)
**Qwen-Audio-3.0-TTS (Flash / Plus)** (Alibaba (Qwen / Tongyi Lab); current; audio/speech; released 2026-07-20) | Alibaba Cloud Model Studio (Singapore / Beijing): `qwen-audio-3.0-tts-flash`; Alibaba Cloud Model Studio: `qwen-audio-3.0-tts-plus` — Flash tier targets real-time use (~300 ms first packet, press); Plus targets quality (throughput ~16 chars/s, press). Languages: ar, zh, en, fr, de, id, it, ja, ko, ms, pt, ru, es, tl, th, vi. Companion qwen-audio-3.0-realtime-plus/-flash and qwen-audio-3.0-asr-flash also exist. Superseded by Qwen-Audio-3.1 (2026-09-23), but as of 2026-09-29 the Model Studio catalog still lists qwen-audio-3.0-tts-plus as its TTS model, and no 3.1 TTS id is published in the international docs.
  - Alibaba Cloud Model Studio (Singapore / Beijing): `qwen-audio-3.0-tts-flash` — https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer (docs: https://www.alibabacloud.com/help/en/model-studio/qwen-tts)
  - Alibaba Cloud Model Studio: `qwen-audio-3.0-tts-plus` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - Pricing source: https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/
Notable capabilities:
  - #1 on Artificial Analysis TTS arena at launch: Qwen-Audio-3.0-TTS-Plus ranked first on the Artificial Analysis Text-to-Speech leaderboard in July 2026 (Elo ~1,236-1,237, just ahead of Speechify Simba 3.2 at ~1,234). It was later overtaken (Eleven v4 was #1 by late Sept 2026). (https://arxiv.org/abs/2607.23938)
  - Controllable, robust multilingual synthesis: 12.5 Hz speech tokenizer plus a five-stage LM + flow-matching training recipe; natural-language instructions and inline tags; 16 languages and 20 Chinese dialect regions; one-pass long-form output up to 3 minutes; voice cloning works from noisy or reverberant references. (https://arxiv.org/abs/2607.23938)


## Alibaba (Qwen)

### Qwen-Audio-3.1-ASR (Flash)
**Qwen-Audio-3.1-ASR (Flash)** (Alibaba (Qwen); current; audio/speech; released 2026-09-23) | Alibaba Cloud Model Studio (streaming): `qwen-audio-3.1-asr-flash-streaming`; Alibaba Cloud Model Studio / QwenCloud (file transcription): `qwen-audio-3.1-asr-flash-filetrans` — Secondary sources report 30 languages + Chinese dialects and ~160 ms latency (unverified). Sibling Qwen-Audio-3.1-ASR-Next adds speaker diarization with timestamps, emotion and sound-event detection (API id not verified). Previous: qwen-audio-3.0-asr-flash; open-weights alternative Qwen3-ASR (see qwen3-asr). Pricing not verified on an official page.
  - Alibaba Cloud Model Studio (streaming): `qwen-audio-3.1-asr-flash-streaming` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - Alibaba Cloud Model Studio / QwenCloud (file transcription): `qwen-audio-3.1-asr-flash-filetrans` (docs: https://www.qwencloud.com/models/qwen-audio-3.1-asr-flash-filetrans)
Notable capabilities:
  - Multilingual + dialect ASR with disfluency cleanup: Improved multilingual and Chinese-dialect recognition that automatically removes filler words and repetitions; launched with up to 95% price cut. (https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/)

### Qwen-Audio-3.1-Realtime (Plus)
**Qwen-Audio-3.1-Realtime (Plus)** (Alibaba (Qwen); current; audio/speech; released 2026-09-23) | ctx 262,144 | QwenCloud (Realtime WebSocket): `qwen-audio-3.1-realtime-plus`; Alibaba Cloud Model Studio (Singapore / Beijing): `qwen-audio-3.1-realtime-plus` — Languages: de, en, es, fr, id, it, ja, ko, pt, ru, zh (Mandarin, Cantonese and 18+ Chinese varieties). Predecessors qwen-audio-3.0-realtime-plus / -flash (July 2026) still listed. Press (MarkTechPost) reports interruption-stop latency 1.116 s vs 0.383 s for GPT-Realtime-2 and higher red-team refusal for GPT-Realtime-2; not verified on an official page. Release date is the announcement date (Qwen X post / Apsara); Model Studio pricing for this id not verified.
  - QwenCloud (Realtime WebSocket): `qwen-audio-3.1-realtime-plus` — wss://maas.qwencloudapi.com/api-ws/v1/realtime?model=qwen-audio-3.1-realtime-plus (docs: https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus)
  - Alibaba Cloud Model Studio (Singapore / Beijing): `qwen-audio-3.1-realtime-plus` — wss://{WorkspaceId}.sg-singapore.maas.aliyuncs.com/api-ws/v1/realtime (docs: https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides)
  - Pricing source: https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus
Notable capabilities:
  - Full-duplex agentic voice ("Think, Act, Speak and Coordinate"): Listens while speaking, decides whether to keep listening, speak, stop or resume; function calling and built-in web search. Task success 82.0% vs 78.4% for the previous version; replies to background speech fell from 73.0% to 13.0% (Full-Duplex-Bench v1.5). (https://arxiv.org/abs/2609.25176)
  - Three turn-taking modes and voice cloning: server_vad, semantic smart_turn and push-to-talk modes; system voices plus cloned custom voices; 16 kHz PCM in, 24 kHz PCM out. (https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides)
  - ~85% price cut at launch: Alibaba cut Realtime prices about 85% with the 3.1 release (TTS ~70%, ASR up to 95%). (https://x.com/Alibaba_Qwen/status/2102687258990026993)

### Qwen-Audio-3.1-TTS-Next
**Qwen-Audio-3.1-TTS-Next** (Alibaba (Qwen); current; audio/speech; released 2026-09-22) | $0.848 in / $1.696 out USD per 1M tokens (China/Beijing region price shown in docs; international price not listed) | Alibaba Cloud Model Studio: `qwen-audio-3.1-tts-next` — Chinese and English only; max 3,000 input characters; output up to 240 s for podcasts, 120 s otherwise. Comparable to ByteDance Seed Audio 1.0 (Jul 2026) and StepAudio 3 Gen. Sibling TTS model Qwen-Audio-3.1-TTS (plain TTS, ~70% cheaper than 3.0) exists but its exact API id was not verified: as of 2026-09-29 the international Model Studio docs (models page, qwen-tts page) list only qwen-audio-3.0-tts-flash / -plus, and neither qwen-audio-3.1-tts-flash/-plus nor an ASR-Next id resolves on QwenCloud (404). Verified 3.1 ASR ids: qwen-audio-3.1-asr-flash(-streaming/-filetrans).
  - Alibaba Cloud Model Studio: `qwen-audio-3.1-tts-next` (docs: https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next)
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next
Notable capabilities:
  - One-pass speech + sound effects + ambience: 'AudioGen' model (LM + diffusion) that generates complete audio - speech, multi-speaker dialogue, podcasts, sound effects and ambient soundscapes - in a single pass from text, timestamps and up to 3 reference clips. (https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next)

### Qwen3.8-LiveTranslate (Flash Realtime)
**Qwen3.8-LiveTranslate (Flash Realtime)** (Alibaba (Qwen); current; audio/speech; released 2026-09-19) | ctx 53,248 | QwenCloud (Realtime WebSocket): `qwen3.8-livetranslate-flash-realtime`; Alibaba Cloud Model Studio: `qwen3.8-livetranslate-flash-realtime` — Understands 60 languages and speaks 29 (the rest get text-only translation). Thinker-talker hybrid MoE on the Qwen-Omni stack (press). API-only, no open weights and no announced timeline for them. MindStudio's hands-on found short sentences fine but weak end-of-turn detection, so developers need their own turn-taking logic. Announced on X 2026-09-19 (294k views by 2026-09-29), shortly before Apsara 2026.
  - QwenCloud (Realtime WebSocket): `qwen3.8-livetranslate-flash-realtime` — wss://maas.qwencloudapi.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime (docs: https://www.qwencloud.com/models/qwen3.8-livetranslate-flash-realtime)
  - Alibaba Cloud Model Studio: `qwen3.8-livetranslate-flash-realtime` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - Pricing source: https://www.qwencloud.com/models/qwen3.8-livetranslate-flash-realtime
Notable capabilities:
  - Simultaneous interpretation with lower lag: Streams translated speech and text while the speaker is still talking; average lagging (LAAL) cut from 2.8 s to 2.3 s across 60 languages with a new 'Interleave' architecture. (https://x.com/Alibaba_Qwen/status/2101206705111757253)
  - Multi-speaker diarization with per-speaker voice cloning: Tells speakers apart in multi-party speech and keeps each speaker's own voice in the translated audio; synchronized bilingual on-screen display. (https://x.com/Alibaba_Qwen/status/2101206705111757253)
  - Long-context disambiguation: Uses conversation history to keep names and terminology consistent across a session. (https://x.com/Alibaba_Qwen/status/2101206705111757253)

### Qwen3.8-Omni-Flash
**Qwen3.8-Omni-Flash** (Alibaba (Qwen); current; multimodal; released 2026-09) | ctx 1,000,000 | $0.15 in / $0.47 out per 1M tokens (USD), Singapore/International | Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-omni-flash`; Alibaba Cloud Model Studio (realtime voice/video): `qwen3.8-omni-flash-realtime`; OpenRouter: `qwen/qwen3.8-omni-flash` | Web app: https://chat.qwen.ai — Thinking on by default with adjustable effort. Realtime variant qwen3.8-omni-flash-realtime: $0.93 audio in / $1.87 audio out per 1M tokens (Singapore/Intl pricing page, checked 2026-09-29). For dedicated hosted voice agents Alibaba also offers qwen-audio-3.1-realtime-plus (see qwen-audio-3-1-realtime). Release day not verified (OpenRouter listing 2026-09-21).
  - Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-omni-flash` — https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 (docs: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash)
  - Alibaba Cloud Model Studio (realtime voice/video): `qwen3.8-omni-flash-realtime` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - OpenRouter: `qwen/qwen3.8-omni-flash` — https://openrouter.ai/qwen/qwen3.8-omni-flash
  - Web app — https://chat.qwen.ai
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/model-pricing
Notable capabilities:
  - Audio + video understanding with 1M context: Text, image, audio and video in, text out; 113 input languages/dialects for audio. (https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash)
  - Spatial (multichannel) audio input: Accepts multichannel/spatial audio via use_multichannel in Chat Completions. (https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash)
  - Realtime speech-to-speech sibling: qwen3.8-omni-flash-realtime handles live audio/video conversation; for non-realtime audio output Alibaba points to qwen3.5-omni-plus. (https://www.alibabacloud.com/help/en/model-studio/models)

### Qwen3.8-27B
**Qwen3.8-27B** (Alibaba (Qwen); current; multimodal; released 2026-08-05; open weights) | ctx 262,144 | OpenRouter: `qwen/qwen3.8-27b`; OpenRouter (free tier): `qwen/qwen3.8-27b:free` | Hugging Face: https://huggingface.co/Qwen/Qwen3.8-27B; Hugging Face (FP8): https://huggingface.co/Qwen/Qwen3.8-27B-FP8; Web app: https://chat.qwen.ai — Best Apache-2.0 Qwen for self-hosting; also the go-to open Qwen VL model (Qwen3-VL successor). First-party hosted API 'coming soon' on Qwen Cloud at time of check. Pricing not verified (no first-party price).
  - Hugging Face — https://huggingface.co/Qwen/Qwen3.8-27B
  - Hugging Face (FP8) — https://huggingface.co/Qwen/Qwen3.8-27B-FP8
  - OpenRouter: `qwen/qwen3.8-27b` — https://openrouter.ai/qwen/qwen3.8-27b
  - OpenRouter (free tier): `qwen/qwen3.8-27b:free` — https://openrouter.ai/qwen/qwen3.8-27b
  - Web app — https://chat.qwen.ai
Notable capabilities:
  - Dense open VLM with agentic focus: 27B dense native vision-language model (images and hour-scale video) tuned for coding and long-horizon agent tasks, Apache-2.0. (https://huggingface.co/Qwen/Qwen3.8-27B)
  - Thinking control: Thinking on by default, can be disabled per request; reasoning_effort and preserve_thinking supported. (https://huggingface.co/Qwen/Qwen3.8-27B)
  - Extensible to 1M context: 262,144 tokens native, extensible up to 1,000,000. (https://huggingface.co/Qwen/Qwen3.8-27B)

### Qwen3.8-Flash
**Qwen3.8-Flash** (Alibaba (Qwen); current; reasoning-llm; released 2026-08; open weights) | ctx 1,000,000 | $0.15 in / $0.47 out per 1M tokens (USD), Singapore/International region, input up to 1M | Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-flash`; OpenRouter: `qwen/qwen3.8-flash` | Hugging Face (Qwen3.8-Flash-Next, base of the API model): https://huggingface.co/Qwen/Qwen3.8-Flash-Next; Web app: https://chat.qwen.ai — Low-cost default in Model Studio (maps to 'GPT-5.4-mini / Haiku 4.5' tier per Alibaba). Max output not verified. Release day not verified (OpenRouter listing 2026-08-26).
  - Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-flash` — https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 (docs: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash)
  - OpenRouter: `qwen/qwen3.8-flash` — https://openrouter.ai/qwen/qwen3.8-flash
  - Hugging Face (Qwen3.8-Flash-Next, base of the API model) — https://huggingface.co/Qwen/Qwen3.8-Flash-Next
  - Web app — https://chat.qwen.ai
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/model-pricing
Notable capabilities:
  - Preview of the Qwen4 architecture: Built on Qwen3.8-Flash-Next, an experimental preview of the architecture that will underpin Qwen4 (Gated DeltaNet + Qwen Sparse Attention, Gated Residual, N-gram Embedding). (https://huggingface.co/Qwen/Qwen3.8-Flash-Next)
  - Block-level sparse attention (QSA): Qwen Sparse Attention selects micro-blocks rather than tokens, cutting long-context latency for agentic workloads. (https://huggingface.co/Qwen/Qwen3.8-Flash-Next)
  - OpenAI + Anthropic protocol compatibility: Works directly with Claude Code and Codex; 1M context, image/video understanding, desktop-app operation. (https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash)

### Qwen3.8-Max
**Qwen3.8-Max** (Alibaba (Qwen); current; reasoning-llm; released 2026-08; open weights) | ctx 1,000,000 | $2 in / $6 out per 1M tokens (USD), Singapore/International region, input up to 1M; Beijing/Global regions 1.65/4.951 | Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-max`; Alibaba Cloud Model Studio (US Virginia): `qwen3.8-max`; OpenRouter: `qwen/qwen3.8-max-0902`; OpenRouter (open-weight base): `qwen/qwen3.8-2.4t-a95b` | Hugging Face: https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B; Web app: https://chat.qwen.ai — Alibaba's top model. Apsara 2026 (2026-09-22): Alibaba says an updated Qwen3.8-Max went through 33 automated self-improvement cycles, raising its Artificial Analysis score from 40 to 45 (company claim, https://www.alibabacloud.com/en/press-room/alibaba-unveils-roadmap-on-full-stack-ai-strategy). Snapshot qwen3.8-max-0902; fast tier qwen3.8-max-prime (OpenRouter qwen/qwen3.8-max-prime, Beijing 3.301/9.902). Singapore endpoint needs your WorkspaceId (old dashscope-intl domain is being migrated). Also sold via Qwen Cloud (qwencloud.com). Release day not verified (weights on HF 2026-08-08). Knowledge cutoff not published.
  - Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.8-max` — https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 (docs: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max)
  - Alibaba Cloud Model Studio (US Virginia): `qwen3.8-max` — https://dashscope-us.aliyuncs.com/compatible-mode/v1 (docs: https://www.alibabacloud.com/help/en/model-studio/compatibility-of-openai-with-dashscope)
  - OpenRouter: `qwen/qwen3.8-max-0902` — https://openrouter.ai/qwen/qwen3.8-max-0902
  - OpenRouter (open-weight base): `qwen/qwen3.8-2.4t-a95b` — https://openrouter.ai/qwen/qwen3.8-2.4t-a95b
  - Hugging Face — https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
  - Web app — https://chat.qwen.ai
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max
Notable capabilities:
  - First open-weight Qwen-Max-class model: Qwen3.8 brings a Max-class model to open release for the first time (Qwen3.8-2.4T-A95B, 2.4T total / 95B active MoE); the API version adds vision input, non-thinking mode, 1M context and built-in tools. (https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)
  - Multi-day autonomous coding: Alibaba markets it as able to code autonomously for over ten days to deliver complete projects, with closed-loop planning and iteration. (https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max)
  - Native vision in the agent loop: Image and video understanding used throughout planning, execution and verification; parses ultra-long documents and long videos. (https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max)
  - Tunable and preserved thinking: reasoning_effort controls depth; preserve_thinking keeps reasoning context from earlier turns. (https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B)

### Qwen-Image-3.0 (Pro)
**Qwen-Image-3.0 (Pro)** (Alibaba (Qwen); current; image-gen; released 2026-07-21) | Alibaba Cloud Model Studio: `qwen-image-3.0-pro`; Alibaba Cloud Model Studio: `qwen-image-3.0` | Hugging Face (open sibling Qwen-Image-2.1, research license): https://huggingface.co/Qwen/Qwen-Image-2.1; Web app: https://chat.qwen.ai — Released 2026-07-21 (invite-only for two weeks, opened to Qwen app users 2026-08-05, per press). Open-weight alternative: Qwen-Image-2.1 (7B DiT, 2026-09-14, qwen-research license). API endpoint path not verified here - see docs.
  - Alibaba Cloud Model Studio: `qwen-image-3.0-pro` (docs: https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro)
  - Alibaba Cloud Model Studio: `qwen-image-3.0` (docs: https://www.alibabacloud.com/help/en/model-studio/model-pricing)
  - Hugging Face (open sibling Qwen-Image-2.1, research license) — https://huggingface.co/Qwen/Qwen-Image-2.1
  - Web app — https://chat.qwen.ai
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/model-pricing
Notable capabilities:
  - Dense single-pass layouts: Prompts up to ~4.5K tokens; generates newspapers, storyboards, menus, exam papers and images-within-images in one pass. (https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro)
  - Tiny, multilingual text rendering: Legible text down to ~10px, native rendering of 12 languages and multiple fonts, realistic UI simulation (web pages, games, livestreams). (https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro)
  - Closed release (break from open Qwen-Image) (found after launch): Shipped without weights, benchmarks or model card, unlike earlier open Qwen-Image releases. (https://www.unite.ai/alibaba-launches-qwen-image-3-0-without-benchmarks-or-weights/)

### Qwen3.7-Plus
**Qwen3.7-Plus** (Alibaba (Qwen); current; reasoning-llm; released 2026-05-26) | ctx 1,000,000 | $0.4 in / $1.6 out per 1M tokens (USD), Singapore/International, input up to 256K (list price; limited-time 20% off). 256K-1M input: 1.2 / 4.8 | Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.7-plus`; OpenRouter: `qwen/qwen3.7-plus` | Web app: https://chat.qwen.ai — Alias of snapshot qwen3.7-plus-2026-05-26 (release date taken from the snapshot name). Thinking and non-thinking modes. Max output not verified.
  - Alibaba Cloud Model Studio (DashScope, Singapore/Intl): `qwen3.7-plus` — https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1 (docs: https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus)
  - OpenRouter: `qwen/qwen3.7-plus` — https://openrouter.ai/qwen/qwen3.7-plus
  - Web app — https://chat.qwen.ai
  - Pricing source: https://www.alibabacloud.com/help/en/model-studio/model-pricing
Notable capabilities:
  - Multimodal hybrid GUI agent: Perceives real-world scenes, reads screens and operates GUIs, generates code from visual references and navigates mobile apps end to end. (https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus)
  - Recommended balanced coding model (found after launch): Alibaba's recommended model for coding tools: full tool calling, built-in tools and 1M context at mid-tier price. (https://www.alibabacloud.com/help/en/model-studio/text-generation-model)

### Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner
**Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner** (Alibaba (Qwen); current; audio/speech; released 2026-01-29; open weights) | Hugging Face: `Qwen/Qwen3-ASR-1.7B`; Hugging Face: `Qwen/Qwen3-ASR-0.6B`; Hugging Face: `Qwen/Qwen3-ForcedAligner-0.6B` | GitHub: https://github.com/QwenLM/Qwen3-ASR — Native Transformers (-hf repos) support added 2026-06-26. Hosted ASR is now Qwen-Audio-3.x-ASR (see qwen-audio-3-1-asr).
  - Hugging Face: `Qwen/Qwen3-ASR-1.7B` — https://huggingface.co/Qwen/Qwen3-ASR-1.7B
  - Hugging Face: `Qwen/Qwen3-ASR-0.6B` — https://huggingface.co/Qwen/Qwen3-ASR-0.6B
  - Hugging Face: `Qwen/Qwen3-ForcedAligner-0.6B` — https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B
  - GitHub — https://github.com/QwenLM/Qwen3-ASR
Notable capabilities:
  - 52 languages/dialects incl. singing and music: Language ID + ASR for 30 languages and 22 Chinese dialects, robust on songs/music; built on Qwen3-Omni audio understanding; vLLM batch and streaming inference, timestamp prediction via ForcedAligner. (https://github.com/QwenLM/Qwen3-ASR)
  - Beats Whisper-large-v3 on Chinese: Self-reported WER e.g. AISHELL-2 2.71 vs 5.06 (Whisper-large-v3); Cantonese CV-yue 7.57 vs 11.36 (GPT-4o-Transcribe). (https://github.com/QwenLM/Qwen3-ASR)

### Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)
**Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)** (Alibaba (Qwen); current; audio/speech; released 2026-01-22; open weights) | Hugging Face: `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`; Hugging Face: `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign`; Alibaba Cloud Model Studio: `qwen3-tts-flash`; Alibaba Cloud Model Studio (instruct / voice design / voice clone): `qwen3-tts-instruct-flash` | GitHub: https://github.com/QwenLM/Qwen3-TTS — HF repos: Qwen3-TTS-12Hz-{1.7B,0.6B}-{Base,CustomVoice}, 1.7B-VoiceDesign, Qwen3-TTS-Tokenizer-12Hz. API snapshots: qwen3-tts-flash (=2025-11-27), qwen3-tts-flash-2025-09-18, qwen3-tts-instruct-flash-2026-01-26, qwen3-tts-vd-2026-01-26 (voice design), qwen3-tts-vc-2026-01-22 (voice clone). Superseded in Alibaba's hosted lineup by Qwen-Audio-3.0-TTS (Jul 2026) and Qwen-Audio-3.1-TTS (Sep 2026). API pricing not verified.
  - Hugging Face: `Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice` — https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
  - Hugging Face: `Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign` — https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign
  - GitHub — https://github.com/QwenLM/Qwen3-TTS
  - Alibaba Cloud Model Studio: `qwen3-tts-flash` (docs: https://www.alibabacloud.com/help/en/model-studio/qwen-tts)
  - Alibaba Cloud Model Studio (instruct / voice design / voice clone): `qwen3-tts-instruct-flash` (docs: https://www.alibabacloud.com/help/en/model-studio/qwen-tts)
Notable capabilities:
  - Open-weights voice design and 3-second cloning: Voice design from natural-language descriptions and voice cloning from ~3 s of audio, in 10 languages (zh, en, ja, ko, de, fr, ru, pt, es, it). (https://github.com/QwenLM/Qwen3-TTS)
  - 97 ms streaming latency: 12 Hz multi-codebook tokenizer; first audio packet after a single input character, end-to-end latency as low as 97 ms; one model for streaming and non-streaming. (https://arxiv.org/abs/2601.15621)


## Alibaba (Tongyi Lab / FunAudioLLM)

### Fun-CosyVoice3 0.5B (2512) + Fun-ASR-Nano + Fun-Audio-Chat-8B
**Fun-CosyVoice3 0.5B (2512) + Fun-ASR-Nano + Fun-Audio-Chat-8B** (Alibaba (Tongyi Lab / FunAudioLLM); current; audio/speech; released 2025-12-11; open weights) | Hugging Face: `FunAudioLLM/Fun-CosyVoice3-0.5B-2512`; Hugging Face (ASR, 800M): `FunAudioLLM/Fun-ASR-Nano-2512`; Hugging Face (speech chat, 8B): `FunAudioLLM/Fun-Audio-Chat-8B` | GitHub: https://github.com/QwenAudio/CosyVoice — HF repo creation dates: CosyVoice3-0.5B-2512 2025-12-11, Fun-ASR-Nano-2512 2025-12-15, Fun-Audio-Chat-8B 2025-12-23. CosyVoice3-0.5B had ~197k downloads in the month to 2026-09-29, one of the most-used open TTS checkpoints. Papers: CosyVoice 3 arXiv 2505.17589, FunAudio-ASR arXiv 2509.12508, Fun-Audio-Chat arXiv 2512.20156. GitHub repo moved from FunAudioLLM/CosyVoice to QwenAudio/CosyVoice. The same Tongyi group's hosted successors are the Qwen-Audio 3.x API models.
  - Hugging Face: `FunAudioLLM/Fun-CosyVoice3-0.5B-2512` — https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
  - Hugging Face (ASR, 800M): `FunAudioLLM/Fun-ASR-Nano-2512` — https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512
  - Hugging Face (speech chat, 8B): `FunAudioLLM/Fun-Audio-Chat-8B` — https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B
  - GitHub — https://github.com/QwenAudio/CosyVoice
Notable capabilities:
  - Small open multilingual zero-shot TTS: 0.5B model with 9 languages (zh, en, ja, ko, de, es, fr, it, ru) and 18+ Chinese dialects/accents; RL variant reports 0.81% CER / 77.4% speaker similarity (zh) and 1.68% WER / 69.5% similarity (en) on its eval set. (https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512)
  - Compact far-field ASR (Fun-ASR-Nano, 800M): zh/en/ja plus 7 Chinese dialect groups and 26 accents; WER 1.80% AIShell1, 1.76% LibriSpeech-clean; tuned for noisy far-field audio and lyrics over music. MLT-Nano variant covers 31 languages. (https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512)
  - Open 8B speech chat model with function calling (Fun-Audio-Chat): Half-duplex speech-to-speech/speech-to-text LLM (zh/en) with dual-resolution speech representations (5 Hz backbone + 25 Hz head, about 50% less compute); spoken QA, speech function calling, voice empathy. (https://arxiv.org/abs/2512.20156)


## Amazon

### Amazon Nova 2 Lite
**Amazon Nova 2 Lite** (Amazon; current; reasoning-llm; released 2025-12-02) | ctx 1,000,000 | $0.3 in / $2.5 out per 1M tokens (USD) on OpenRouter; Bedrock on-demand price not verified | AWS Bedrock: `amazon.nova-2-lite-v1:0`; OpenRouter: `amazon/nova-2-lite-v1` — Amazon's current GA general model. Nova 2 Pro and Nova 2 Omni were preview-only (Nova Forge) at last check; no Bedrock ids verified.
  - AWS Bedrock: `amazon.nova-2-lite-v1:0` — https://bedrock-runtime.{region}.amazonaws.com (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html)
  - OpenRouter: `amazon/nova-2-lite-v1` — https://openrouter.ai/amazon/nova-2-lite-v1
  - Pricing source: https://openrouter.ai/amazon/nova-2-lite-v1
Notable capabilities:
  - Adjustable extended thinking + 1M context: Nova 2 generation adds adjustable extended thinking and a 1M-token context for text/image/video input. (https://www.aboutamazon.com/news/aws/aws-agentic-ai-amazon-bedrock-nova-models)
  - Built-in code interpreter, web grounding, remote MCP: Nova 2 models support built-in tools (code interpreter, web grounding) and remote MCP tools on Bedrock. (https://www.aboutamazon.com/news/aws/aws-agentic-ai-amazon-bedrock-nova-models)

### Amazon Nova 2 Sonic
**Amazon Nova 2 Sonic** (Amazon; current; audio/speech; released 2025-12-02) | ctx 1,000,000 | AWS Bedrock: `amazon.nova-2-sonic-v1:0` — Technical report (Amazon Nova 2, Dec 2025, https://www.amazon.science/publications/amazon-nova-2-multimodal-reasoning-and-generation-models): Big Bench Audio 87.0 (Artificial Analysis) vs GPT-Realtime (Aug 2025) 83.0 and Gemini 2.5 Flash Live 71.0; BFCL subset 74.5; ComplexFunction 65.2; Common Voice avg WER 6.5 vs 8.4 (GPT-Realtime) across 7 languages; human-preference win rate vs GPT-Realtime above 50% for 6 of 8 voices (e.g. 68.4% Spanish) but 42.4% Hindi and 26.3% Portuguese; vs Gemini 2.5 Flash Live 47.5-77.9%. Comparisons are against 2025 competitors. Successor to Nova Sonic (amazon.nova-sonic-v1:0, Apr 2025). Bedrock only, In-Region in us-east-1, us-west-2, eu-north-1, ap-northeast-1 (no cross-region inference); Standard tier only. Lifecycle Active, EOL no sooner than 2026-12-02. No newer Nova Sonic found as of 2026-09-29; per July 2026 reports Nova 2 Sonic is among the Nova models Amazon keeps developing after its Nova wind-down. Prices from secondary source (AWS Nova pricing page does not list per-token rates).
  - AWS Bedrock: `amazon.nova-2-sonic-v1:0` — https://bedrock-runtime.{region}.amazonaws.com (InvokeModelWithBidirectionalStream) (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html)
  - Pricing source: https://www.deeplearning.ai/the-batch/nova-2-family-boosts-cost-effective-performance-adds-new-agentic-features
Notable capabilities:
  - Real-time speech-to-speech: Single model for natural real-time voice conversations over a bidirectional streaming API (no separate ASR/TTS pipeline). (https://aws.amazon.com/blogs/aws/introducing-amazon-nova-2-sonic-next-generation-speech-to-speech-model-for-conversational-ai/)
  - 1M-token session context: 1M-token context window and 64K max output listed for long-running voice sessions. (https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html)
  - Polyglot voices and turn-taking control: Same voice speaks multiple languages natively (Portuguese and Hindi added vs Nova Sonic); developers set low/medium/high pause sensitivity. (https://aws.amazon.com/about-aws/whats-new/2025/12/amazon-nova-2-sonic-real-time-conversational-ai)

### Amazon Nova Premier
**Amazon Nova Premier** (Amazon; retired; multimodal; released 2025-10-31) | ctx 1,000,000 | $2.5 in / $12.5 out per 1M tokens (USD) on OpenRouter | AWS Bedrock: `amazon.nova-premier-v1:0`; OpenRouter: `amazon/nova-premier-v1` — Bedrock card shows lifecycle Legacy with EOL date 2026-09-14 (passed); may still be listed. Use Nova 2 Lite instead. Launch date as shown on Bedrock card. Nova Pro/Lite/Micro (v1) and Nova Canvas/Reel (EOL 2026-09-30) are also legacy.
  - AWS Bedrock: `amazon.nova-premier-v1:0` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html)
  - OpenRouter: `amazon/nova-premier-v1` — https://openrouter.ai/amazon/nova-premier-v1
  - Pricing source: https://openrouter.ai/amazon/nova-premier-v1
Notable capabilities:
  - Teacher model for distillation: Positioned for complex reasoning, agentic workflows and as a teacher for Bedrock model distillation into smaller Nova models. (https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html)
  - 1M context multimodal reasoning: 1M-token context over text, image and video with reasoning support - largest first-gen Nova. (https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html)


## Anthropic

### Claude Sonnet 5.5
**Claude Sonnet 5.5** (Anthropic; current; reasoning-llm; released 2026-09-28) | ctx 1,000,000 | $2 in / $10 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-sonnet-5-5`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-sonnet-5-5`; Google Cloud Vertex AI: `claude-sonnet-5-5`; Microsoft Foundry (Azure): `claude-sonnet-5-5`; Claude Platform on AWS: `claude-sonnet-5-5`; OpenRouter: `anthropic/claude-sonnet-5.5` | Web app: https://claude.ai — Best speed/intelligence balance. Adaptive thinking on by default (effort default high); thinking {type: disabled} returns 400, use {type: between_tools} at effort high or below; forced tool_choice any/tool returns 400; non-default temperature/top_p/top_k return 400. Batch $1/$5.
  - Anthropic API (Claude API): `claude-sonnet-5-5` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/sonnet-5-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-sonnet-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-sonnet-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-sonnet-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-sonnet-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-sonnet-5.5` — https://openrouter.ai/anthropic/claude-sonnet-5.5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Opus-level knowledge work at Sonnet price: Scores nearly level with Opus 5.5 on GDPval-AA (1844 vs 1846 Elo), at $2/$10 per MTok. (https://www.anthropic.com/claude-sonnet-5-5)
  - Large agentic-coding jump: Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5, and up to 30% lower cost per task. (https://www.anthropic.com/claude-sonnet-5-5)
  - [FIRST] Beat Pokemon Red from screenshots: Anthropic says it is the first Sonnet model to finish Pokemon Red using only screenshots. (https://www.anthropic.com/claude-sonnet-5-5)
  - [FIRST] between_tools thinking mode: New thinking type that turns off up-front thinking while still reasoning between tool calls; it replaces thinking: disabled. (https://platform.claude.com/docs/en/models/sonnet-5-5/overview)
  - Token efficiency: A Balyasny test used 121k tokens per task, versus 497k for Sonnet 5. (https://www.anthropic.com/claude-sonnet-5-5)

### Claude Opus 5.5
**Claude Opus 5.5** (Anthropic; current; reasoning-llm; released 2026-09-22) | ctx 1,000,000 | $4 in / $20 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-5-5`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-5-5`; Google Cloud Vertex AI: `claude-opus-5-5`; Microsoft Foundry (Azure): `claude-opus-5-5`; Claude Platform on AWS: `claude-opus-5-5`; OpenRouter: `anthropic/claude-opus-5.5` | Web app: https://claude.ai — Anthropic's recommended default model. Thinking always on (cannot be disabled); effort default is medium (set explicitly); forced tool_choice any/tool returns 400; computer use only via computer_toolset_20260801 on Claude API/Google Cloud. Fast mode (Claude API only) $8/$40. Batch $2/$10; up to 300K output on Batch with output-300k-2026-03-24 beta.
  - Anthropic API (Claude API): `claude-opus-5-5` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-5-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-opus-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-5-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-5.5` — https://openrouter.ai/anthropic/claude-opus-5.5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Top agentic coding at lower cost: Anthropic reports 66.4% on Terminal-Bench 4.0, ahead of GPT-6 Astra at roughly 40% of the cost; an early tester finished a 680k-line code migration in under a day. (https://www.anthropic.com/claude-opus-5-5)
  - Knowledge-work lead (GDPval-AA): Launch claim of 1846 Elo on GDPval-AA v2.1, above both Claude Fable 5.1 (1735) and Claude Opus 5 (1708). (https://www.anthropic.com/claude-opus-5-5)
  - Cheaper, faster Opus: About 40% cheaper than Opus 5 on typical workloads ($4/$20 per MTok, cache reads $0.20) and about 30% faster output at default settings. (https://www.anthropic.com/claude-opus-5-5)
  - [FIRST] Opus with Fable-level safeguards: Anthropic says it is the first Opus model whose safeguards match Claude Fable 5.1 on cyber, bio and distillation (refusal categories include bio and reasoning_extraction). (https://www.anthropic.com/claude-opus-5-5)
  - Thinking that cannot be disabled: Adaptive thinking is always on and effort is the only control (default medium). Text between tool calls comes back as progress-update thinking blocks. (https://platform.claude.com/docs/en/models/opus-5-5/overview)

### Claude Fable 5.1
**Claude Fable 5.1** (Anthropic; current; reasoning-llm; released 2026-09-01) | ctx 1,000,000 | $10 in / $50 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-fable-5-1`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-fable-5-1`; Google Cloud Vertex AI: `claude-fable-5-1`; Microsoft Foundry (Azure): `claude-fable-5-1`; Claude Platform on AWS: `claude-fable-5-1`; OpenRouter: `anthropic/claude-fable-5.1` | Web app: https://claude.ai — Anthropic's most capable widely released model; thinking always on (adaptive, effort low..max, default high); forced tool_choice any/tool returns 400; no prefill; 30-day data retention required (no ZDR unless authorized); no Priority Tier. Batch $5/$25.
  - Anthropic API (Claude API): `claude-fable-5-1` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/fable-5-1/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-fable-5-1` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-fable-5-1` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-fable-5-1` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-fable-5-1` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-fable-5.1` — https://openrouter.ai/anthropic/claude-fable-5.1
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Scientific discovery (protein design): In Anthropic's launch examples, its protein designs reached about 10x higher binding affinity than competition winners, with a hit rate near 50%. (https://www.anthropic.com/claude-fable-and-mythos-5-1)
  - Rare-bug hunting: Anthropic reports it found the cause of a one-in-a-million crash that engineers had not explained for years. (https://www.anthropic.com/claude-fable-and-mythos-5-1)
  - Top CursorBench score: Scored 73.4% on CursorBench 3.2.0 at max effort, which Cursor called the most capable model it had run. (https://www.anthropic.com/claude-fable-and-mythos-5-1)
  - [FIRST] Preserved thinking and content provenance: Thinking blocks are bound to the model and the conversation, and editing earlier turns invalidates them. Also adds per-message effort, turn-scoped system messages and content provenance. (https://platform.claude.com/docs/en/models/fable-5-1/overview)
  - Cheaper cache reads: Cache reads cost $0.25/MTok (0.025x input). Anthropic cites up to 45% savings on agentic work compared with Fable 5. (https://platform.claude.com/docs/en/about-claude/pricing)

### Claude Mythos 5.1
**Claude Mythos 5.1** (Anthropic; current; reasoning-llm; released 2026-09-01) | ctx 1,000,000 | $10 in / $50 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-mythos-5-1` — Invitation-only (Project Glasswing, defensive cybersecurity). Same capabilities/pricing as Claude Fable 5.1; not offered on Claude Platform on AWS. Cloud ids not listed publicly; contact Anthropic/AWS/Google account team. Successor to claude-mythos-5 and claude-mythos-preview (deprecated 2026-06-09).
  - Anthropic API (Claude API): `claude-mythos-5-1` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/mythos-5-1/overview)
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Frontier cyber-defense model: Offered only to Project Glasswing participants for defensive cybersecurity. It has the same capabilities as Fable 5.1, with safeguards that depend on the access program. (https://www.anthropic.com/claude-fable-and-mythos-5-1)
  - Scientific discovery: Shares Fable 5.1's launch results, e.g. protein designs with about 10x higher binding affinity than competition winners. (https://www.anthropic.com/claude-fable-and-mythos-5-1)

### Claude Haiku 4.5
**Claude Haiku 4.5** (Anthropic; current; reasoning-llm; released 2025-10-15) | ctx 200,000 | $1 in / $5 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-haiku-4-5-20251001`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-haiku-4-5`; AWS Bedrock (InvokeModel): `anthropic.claude-haiku-4-5-20251001-v1:0`; Google Cloud Vertex AI: `claude-haiku-4-5@20251001`; Microsoft Foundry (Azure): `claude-haiku-4-5`; Claude Platform on AWS: `claude-haiku-4-5`; OpenRouter: `anthropic/claude-haiku-4.5` | Web app: https://claude.ai — Fastest/cheapest current Claude. Snapshot claude-haiku-4-5-20251001 (alias claude-haiku-4-5). Uses extended thinking (thinking type enabled + budget_tokens), no effort parameter. Training data cutoff Jul 2025. Retirement not sooner than 2026-10-15. Batch $0.50/$2.50.
  - Anthropic API (Claude API): `claude-haiku-4-5-20251001` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/haiku-4-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-haiku-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - AWS Bedrock (InvokeModel): `anthropic.claude-haiku-4-5-20251001-v1:0` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy)
  - Google Cloud Vertex AI: `claude-haiku-4-5@20251001` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-haiku-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-haiku-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-haiku-4.5` — https://openrouter.ai/anthropic/claude-haiku-4.5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Sonnet-4-class coding at Haiku price: 73.3% on SWE-bench Verified, roughly matching Sonnet 4 at one-third the cost and over 2x the speed. (https://www.anthropic.com/news/claude-haiku-4-5)
  - Sub-agent workhorse: Reaches about 90% of Sonnet 4.5 on Augment's agentic eval; Anthropic positions it for multi-agent orchestration. (https://www.anthropic.com/news/claude-haiku-4-5)
  - [FIRST] Haiku with extended thinking and computer use: First Haiku model with extended thinking; it also surpasses Sonnet 4 on some computer-use tasks. (https://www.anthropic.com/news/claude-haiku-4-5)

### Claude Opus 5
**Claude Opus 5** (Anthropic; legacy; reasoning-llm; released 2026-07-24) | ctx 1,000,000 | $5 in / $25 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-5`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-5`; Google Cloud Vertex AI: `claude-opus-5`; Microsoft Foundry (Azure): `claude-opus-5`; Claude Platform on AWS: `claude-opus-5`; OpenRouter: `anthropic/claude-opus-5` | Web app: https://claude.ai — Superseded by claude-opus-5-5 (cheaper). Thinking on by default; {type: disabled} allowed only at effort high or below. Fast mode $10/$50 (Claude API only). Retirement not sooner than 2027-07-24.
  - Anthropic API (Claude API): `claude-opus-5` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-opus-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-5` — https://openrouter.ai/anthropic/claude-opus-5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Novel problem solving (ARC-AGI 3): Anthropic says it scored about 3x as high as competing models on ARC-AGI 3. (https://www.anthropic.com/news/claude-opus-5)
  - Near-Fable coding at half the price: Launch claim: more than doubles Opus 4.8 on Frontier-Bench and beats Fable 5 on OSWorld 2.0 at about a third of the cost. (https://www.anthropic.com/news/claude-opus-5)
  - Self-built tooling: In one demo it wrote its own vision pipeline to solve a FreeCAD reconstruction task. (https://www.anthropic.com/news/claude-opus-5)
  - [FIRST] Thinking on by default: First Opus where omitting the thinking parameter runs adaptive thinking. (https://platform.claude.com/docs/en/models/opus-5/overview)

### Claude Sonnet 5
**Claude Sonnet 5** (Anthropic; legacy; reasoning-llm; released 2026-06-30) | ctx 1,000,000 | $2 in / $10 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-sonnet-5`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-sonnet-5`; Google Cloud Vertex AI: `claude-sonnet-5`; Microsoft Foundry (Azure): `claude-sonnet-5`; Claude Platform on AWS: `claude-sonnet-5`; OpenRouter: `anthropic/claude-sonnet-5` | Web app: https://claude.ai — Superseded by claude-sonnet-5-5. $2/$10 introductory price became permanent (planned increase to $3/$15 cancelled). Retirement not sooner than 2027-06-30.
  - Anthropic API (Claude API): `claude-sonnet-5` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/sonnet-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-sonnet-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-sonnet-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-sonnet-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-sonnet-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-sonnet-5` — https://openrouter.ai/anthropic/claude-sonnet-5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Opus 4.8-level quality at Sonnet price: Anthropic positioned it at parity with Opus 4.8 on many tasks at $2/$10 per MTok. (https://www.anthropic.com/news/claude-sonnet-5)
  - Self-verification: Early testers reported it checks its own output without prompting and finishes multi-step workflows where earlier Sonnets stopped short. (https://www.anthropic.com/news/claude-sonnet-5)
  - Cyber safeguards on by default: It launched with deliberately reduced exploit-development capability and with cyber safeguards enabled. (https://www.anthropic.com/news/claude-sonnet-5)

### Claude Fable 5
**Claude Fable 5** (Anthropic; legacy; reasoning-llm; released 2026-06-09) | ctx 1,000,000 | $10 in / $50 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-fable-5`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-fable-5`; Google Cloud Vertex AI: `claude-fable-5`; Microsoft Foundry (Azure): `claude-fable-5`; Claude Platform on AWS: `claude-fable-5`; OpenRouter: `anthropic/claude-fable-5` | Web app: https://claude.ai — Superseded by claude-fable-5-1 (same price, cheaper cache reads). Still served; retirement not sooner than 2027-06-09. Sibling claude-mythos-5 (Project Glasswing only, no safety classifiers).
  - Anthropic API (Claude API): `claude-fable-5` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/fable-5/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-fable-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-fable-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-fable-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-fable-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-fable-5` — https://openrouter.ai/anthropic/claude-fable-5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Strongest cybersecurity capabilities (Mythos 5): Anthropic called the Fable 5 / Mythos 5 generation the 'strongest cybersecurity capabilities of any model in the world'. Mythos 5 runs without safety classifiers for Glasswing defenders. (https://www.anthropic.com/news/claude-fable-5-mythos-5)
  - [FIRST] Rebuild web apps from screenshots: Anthropic claims it was the first model to rebuild a web app's source code from screenshots alone. It also completed Pokemon FireRed using vision only. (https://www.anthropic.com/news/claude-fable-5-mythos-5)
  - Massive code migrations: Stripe reported a 50-million-line migration done in one day instead of about two months. (https://www.anthropic.com/news/claude-fable-5-mythos-5)
  - Novel scientific hypotheses: In blind comparisons, scientists preferred its molecular-biology hypotheses about 80% of the time over Opus-class models. (https://www.anthropic.com/news/claude-fable-5-mythos-5)
  - [FIRST] Refusal stop reason with fallbacks: Safety classifiers can decline a request with stop_reason 'refusal'. A server-side fallbacks parameter retries on another Claude model. (https://platform.claude.com/docs/en/models/fable-5/overview)

### Claude Opus 4.8
**Claude Opus 4.8** (Anthropic; legacy; reasoning-llm; released 2026-05-28) | ctx 1,000,000 | $5 in / $25 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-4-8`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-4-8`; Google Cloud Vertex AI: `claude-opus-4-8`; Microsoft Foundry (Azure): `claude-opus-4-8`; Claude Platform on AWS: `claude-opus-4-8`; OpenRouter: `anthropic/claude-opus-4.8` | Web app: https://claude.ai — Last Opus 4.x. Adaptive thinking only (omit thinking = no thinking); sampling params and budget_tokens removed. Fast mode $10/$50 (Claude API only). Retirement not sooner than 2027-05-28.
  - Anthropic API (Claude API): `claude-opus-4-8` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-4-8/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-4-8` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-opus-4-8` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-4-8` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-4-8` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-4.8` — https://openrouter.ai/anthropic/claude-opus-4.8
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Code honesty: About 4x less likely than Opus 4.7 to let flaws in its own code pass without comment. (https://www.anthropic.com/news/claude-opus-4-8)
  - Browser agents: Scored 84% on Online-Mind2Web, ahead of Opus 4.7 and GPT-5.5. (https://www.anthropic.com/news/claude-opus-4-8)
  - [FIRST] Legal agent benchmark: Anthropic says it was the first model to exceed 10% on the Legal Agent Benchmark all-pass standard. (https://www.anthropic.com/news/claude-opus-4-8)
  - Cheaper fast mode: Fast mode runs up to 2.5x faster, at a lower premium than earlier fast mode. (https://www.anthropic.com/news/claude-opus-4-8)

### Claude Opus 4.7
**Claude Opus 4.7** (Anthropic; legacy; reasoning-llm; released 2026-04-16) | ctx 1,000,000 | $5 in / $25 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-4-7`; AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-4-7`; Google Cloud Vertex AI: `claude-opus-4-7`; Microsoft Foundry (Azure): `claude-opus-4-7`; Claude Platform on AWS: `claude-opus-4-7`; OpenRouter: `anthropic/claude-opus-4.7` | Web app: https://claude.ai — Introduced the newer tokenizer (~30% more tokens per text) and xhigh effort. Adaptive thinking only. Retirement not sooner than 2027-04-16.
  - Anthropic API (Claude API): `claude-opus-4-7` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-4-7/overview)
  - AWS Bedrock (Messages API / Mantle): `anthropic.claude-opus-4-7` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock)
  - Google Cloud Vertex AI: `claude-opus-4-7` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-4-7` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-4-7` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-4.7` — https://openrouter.ai/anthropic/claude-opus-4.7
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - [FIRST] High-resolution vision: Accepts images up to 2,576 px on the long edge (~3.75 MP), more than 3x prior Claude models. Scored 98.5% on XBOW visual acuity versus 54.5% for Opus 4.6. (https://www.anthropic.com/news/claude-opus-4-7)
  - [FIRST] xhigh effort level: Introduced the xhigh effort level between high and max. (https://www.anthropic.com/news/claude-opus-4-7)
  - [FIRST] New tokenizer: First model with Anthropic's newer tokenizer (about 30% more tokens for the same text). (https://platform.claude.com/docs/en/about-claude/pricing)
  - Hard coding tasks: Resolved about 3x more production tasks than Opus 4.6 on Rakuten-SWE-Bench. (https://www.anthropic.com/news/claude-opus-4-7)

### Claude Sonnet 4.6
**Claude Sonnet 4.6** (Anthropic; legacy; reasoning-llm; released 2026-02-17) | ctx 1,000,000 | $3 in / $15 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-sonnet-4-6`; AWS Bedrock (InvokeModel): `anthropic.claude-sonnet-4-6`; Google Cloud Vertex AI: `claude-sonnet-4-6`; Microsoft Foundry (Azure): `claude-sonnet-4-6`; Claude Platform on AWS: `claude-sonnet-4-6`; OpenRouter: `anthropic/claude-sonnet-4.6` | Web app: https://claude.ai — Last model on the older tokenizer. Adaptive thinking (budget_tokens deprecated). Training data cutoff Jan 2026. Bedrock via InvokeModel. Retirement not sooner than 2027-02-17.
  - Anthropic API (Claude API): `claude-sonnet-4-6` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/sonnet-4-6/overview)
  - AWS Bedrock (InvokeModel): `anthropic.claude-sonnet-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy)
  - Google Cloud Vertex AI: `claude-sonnet-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-sonnet-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-sonnet-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-sonnet-4.6` — https://openrouter.ai/anthropic/claude-sonnet-4.6
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - Human-level computer use on common tasks: Anthropic cites human-level performance on tasks such as navigating complex spreadsheets and multi-step web forms (OSWorld). (https://www.anthropic.com/news/claude-sonnet-4-6)
  - Beats previous Opus in user preference: Users preferred it to Opus 4.5 59% of the time on coding, citing less overengineering. (https://www.anthropic.com/news/claude-sonnet-4-6)
  - 1M context for Sonnet 4.6: 1M-token context window (beta at launch). (https://www.anthropic.com/news/claude-sonnet-4-6)

### Claude Opus 4.6
**Claude Opus 4.6** (Anthropic; legacy; reasoning-llm; released 2026-02-05) | ctx 1,000,000 | $5 in / $25 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-4-6`; AWS Bedrock (InvokeModel): `anthropic.claude-opus-4-6-v1`; Google Cloud Vertex AI: `claude-opus-4-6`; Microsoft Foundry (Azure): `claude-opus-4-6`; Claude Platform on AWS: `claude-opus-4-6`; OpenRouter: `anthropic/claude-opus-4.6` | Web app: https://claude.ai — First dateless-ID Opus; adaptive thinking (budget_tokens deprecated). Training data cutoff Aug 2025. Bedrock via InvokeModel only. Retirement not sooner than 2027-02-05.
  - Anthropic API (Claude API): `claude-opus-4-6` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-4-6/overview)
  - AWS Bedrock (InvokeModel): `anthropic.claude-opus-4-6-v1` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy)
  - Google Cloud Vertex AI: `claude-opus-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-4-6` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-4.6` — https://openrouter.ai/anthropic/claude-opus-4.6
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - [FIRST] 1M-token context for Opus: First Opus with a 1M-token context window (launched in beta). Scored 76% on MRCR v2 long-context retrieval versus 18.5% for Sonnet 4.5. (https://www.anthropic.com/news/claude-opus-4-6)
  - [FIRST] Adaptive thinking: Introduced adaptive thinking: the model decides when and how much to think, steered by effort. (https://www.anthropic.com/news/claude-opus-4-6)
  - [FIRST] Agent teams: Research preview of multiple Claude instances coordinating in parallel (in Claude Code). (https://www.anthropic.com/news/claude-opus-4-6)
  - Knowledge work (GDPval-AA): About 144 Elo above GPT-5.2 on GDPval-AA. Also led Terminal-Bench 2.0 at launch. (https://www.anthropic.com/news/claude-opus-4-6)

### Claude Opus 4.5
**Claude Opus 4.5** (Anthropic; legacy; reasoning-llm; released 2025-11-24) | ctx 200,000 | $5 in / $25 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-opus-4-5-20251101`; AWS Bedrock (InvokeModel): `anthropic.claude-opus-4-5-20251101-v1:0`; Google Cloud Vertex AI: `claude-opus-4-5@20251101`; Microsoft Foundry (Azure): `claude-opus-4-5`; Claude Platform on AWS: `claude-opus-4-5`; OpenRouter: `anthropic/claude-opus-4.5` | Web app: https://claude.ai — Snapshot claude-opus-4-5-20251101 (alias claude-opus-4-5). Extended thinking (budget_tokens); effort low/medium/high. Training data cutoff Aug 2025. Retirement not sooner than 2026-11-24.
  - Anthropic API (Claude API): `claude-opus-4-5-20251101` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/opus-4-5/overview)
  - AWS Bedrock (InvokeModel): `anthropic.claude-opus-4-5-20251101-v1:0` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy)
  - Google Cloud Vertex AI: `claude-opus-4-5@20251101` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-opus-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-opus-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-opus-4.5` — https://openrouter.ai/anthropic/claude-opus-4.5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - [FIRST] Beat all human candidates on Anthropic's engineering exam: Scored higher than any human candidate on Anthropic's take-home engineering exam within the 2-hour limit. (https://www.anthropic.com/news/claude-opus-4-5)
  - [FIRST] Effort parameter: First model with the effort parameter. At medium effort it matched Sonnet 4.5's best score with 76% fewer output tokens. (https://www.anthropic.com/news/claude-opus-4-5)
  - Prompt-injection robustness: Anthropic claimed it was harder to trick with prompt injection than any other frontier model at the time. (https://www.anthropic.com/news/claude-opus-4-5)
  - Opus price cut: Opus-class pricing dropped to $5/$25 per MTok, from $15/$75. (https://www.anthropic.com/news/claude-opus-4-5)

### Claude Sonnet 4.5
**Claude Sonnet 4.5** (Anthropic; legacy; reasoning-llm; released 2025-09-29) | ctx 200,000 | $3 in / $15 out per 1M tokens (USD); Batch API 50% off | Anthropic API (Claude API): `claude-sonnet-4-5-20250929`; AWS Bedrock (InvokeModel): `anthropic.claude-sonnet-4-5-20250929-v1:0`; Google Cloud Vertex AI: `claude-sonnet-4-5@20250929`; Microsoft Foundry (Azure): `claude-sonnet-4-5`; Claude Platform on AWS: `claude-sonnet-4-5`; OpenRouter: `anthropic/claude-sonnet-4.5` | Web app: https://claude.ai — Snapshot claude-sonnet-4-5-20250929 (alias claude-sonnet-4-5). Extended thinking only. Training data cutoff Jul 2025. Retirement 'not sooner than 2026-09-29' - may be deprecated soon; check the deprecations page.
  - Anthropic API (Claude API): `claude-sonnet-4-5-20250929` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/models/sonnet-4-5/overview)
  - AWS Bedrock (InvokeModel): `anthropic.claude-sonnet-4-5-20250929-v1:0` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy)
  - Google Cloud Vertex AI: `claude-sonnet-4-5@20250929` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai)
  - Microsoft Foundry (Azure): `claude-sonnet-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry)
  - Claude Platform on AWS: `claude-sonnet-4-5` (docs: https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws)
  - OpenRouter: `anthropic/claude-sonnet-4.5` — https://openrouter.ai/anthropic/claude-sonnet-4.5
  - Web app — https://claude.ai
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - 30+ hour autonomous tasks: Anthropic reported it maintained focus for more than 30 hours on complex multi-step tasks. (https://www.anthropic.com/news/claude-sonnet-4-5)
  - SOTA SWE-bench Verified at launch: 77.2% on SWE-bench Verified; billed as 'the best coding model in the world' at release. (https://www.anthropic.com/news/claude-sonnet-4-5)
  - Computer use lead: 61.4% on OSWorld, up from 42.2% for Sonnet 4. (https://www.anthropic.com/news/claude-sonnet-4-5)

### Claude Opus 4.1
**Claude Opus 4.1** (Anthropic; retired; reasoning-llm; released 2025-08-05) | ctx 200,000 | $15 in / $75 out per 1M tokens (USD), Bedrock/Google Cloud may differ | Anthropic API (Claude API): `claude-opus-4-1-20250805`; OpenRouter: `anthropic/claude-opus-4.1` — Retired on the Claude API 2026-08-05 (replacement claude-opus-4-8 / claude-opus-5-5); still available on Amazon Bedrock and Google Cloud per Anthropic pricing page. Cloud ids not re-verified today.
  - Anthropic API (Claude API): `claude-opus-4-1-20250805` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/about-claude/model-deprecations)
  - OpenRouter: `anthropic/claude-opus-4.1` — https://openrouter.ai/anthropic/claude-opus-4.1
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - SOTA SWE-bench Verified (Aug 2025): 74.5% on SWE-bench Verified at launch. (https://www.anthropic.com/news/claude-opus-4-1)
  - Precise multi-file refactoring: GitHub and Rakuten highlighted multi-file refactoring and pinpoint fixes without unnecessary changes. (https://www.anthropic.com/news/claude-opus-4-1)

### Claude Sonnet 4
**Claude Sonnet 4** (Anthropic; retired; reasoning-llm; released 2025-05-22) | ctx 200,000 | $3 in / $15 out per 1M tokens (USD), Bedrock/Google Cloud may differ | Anthropic API (Claude API): `claude-sonnet-4-20250514`; OpenRouter: `anthropic/claude-sonnet-4` — Retired on the Claude API 2026-06-15 (replacement claude-sonnet-4-6 / claude-sonnet-5-5); still available on Amazon Bedrock and Google Cloud per Anthropic pricing page. Cloud ids not re-verified today.
  - Anthropic API (Claude API): `claude-sonnet-4-20250514` — https://api.anthropic.com/v1/messages (docs: https://platform.claude.com/docs/en/about-claude/model-deprecations)
  - OpenRouter: `anthropic/claude-sonnet-4` — https://openrouter.ai/anthropic/claude-sonnet-4
  - Pricing source: https://platform.claude.com/docs/en/about-claude/pricing
Notable capabilities:
  - [FIRST] Extended thinking with tool use: The Claude 4 generation introduced interleaving tool use (e.g. web search) with extended thinking, plus parallel tool calls. (https://www.anthropic.com/news/claude-4)
  - SOTA SWE-bench at launch: 72.7% on SWE-bench Verified; chosen by GitHub to power the Copilot coding agent. (https://www.anthropic.com/news/claude-4)


## AssemblyAI

### AssemblyAI Universal-3.6 Pro Realtime
**AssemblyAI Universal-3.6 Pro Realtime** (AssemblyAI; current; audio/speech; released 2026-09-29) | AssemblyAI API: `universal-3-6-pro`; AssemblyAI Voice Agent API: `(default STT)` — Lineage: Universal-3 Pro Streaming (Mar 2026) -> Universal-3.5 Pro Realtime (2026-06-23) -> 3.6 (2026-09-29). Older streaming ids u3-rt-pro/u3-pro replaced. Voice Agent API ($4.50/hr all-in: STT+LLM+TTS) GA April 2026. AssemblyAI roadmap targets 30+ native languages for the next Universal-3.x in Q4 2026.
  - AssemblyAI API: `universal-3-6-pro` — wss://streaming.assemblyai.com/v3/ws?model=universal-3-6-pro (docs: https://www.assemblyai.com/blog/universal-3-6-pro-realtime)
  - AssemblyAI Voice Agent API: `(default STT)` — wss://agents.assemblyai.com/v1/ws
  - Pricing source: https://www.assemblyai.com/pricing
Notable capabilities:
  - Promptable streaming STT for voice agents: Prompting + keyterms together, real-time diarization, entity-aware endpointing and native code-switching in 32 languages with auto language detection; 5.13% normalized WER (vs 5.80% for 3.5 Pro Realtime), short-response WER 1.45%; median endpoint latency 537 ms. (https://www.assemblyai.com/blog/universal-3-6-pro-realtime)

### AssemblyAI Universal-3.5 Pro (async)
**AssemblyAI Universal-3.5 Pro (async)** (AssemblyAI; current; audio/speech; released 2026-07-07) | AssemblyAI API: `universal-3-pro`; AssemblyAI Dictation API: `(Universal-3.5 Pro + LLM cleanup)` | OpenRouter: https://openrouter.ai/assemblyai/universal-3-5-pro — API id stays `universal-3-pro` (pass in `speech_models`, plural; singular `speech_model` is deprecated). 18 languages; use universal-2 ($0.15/hr, 99+ languages) for broad coverage and legacy features (auto_chapters/summarization fail on 3.5 Pro). Added to OpenRouter 2026-09-22. Launch date 2026-07-07 from AssemblyAI releases collection via search (not opened directly). Streaming sibling: assemblyai-universal-3-6-pro-realtime. Related AssemblyAI products: Voice Agent API (GA April 2026, $4.50/hr all-in) and LLM Gateway (OpenAI-compatible multi-provider LLM API that replaced LeMUR; migration guide at assemblyai.com/docs/llm-gateway/migration-from-lemur; exact rename date not verified).
  - AssemblyAI API: `universal-3-pro` — https://api.assemblyai.com/v2/transcript (docs: https://www.assemblyai.com/docs/getting-started/universal-3-5-pro)
  - OpenRouter — https://openrouter.ai/assemblyai/universal-3-5-pro
  - AssemblyAI Dictation API: `(Universal-3.5 Pro + LLM cleanup)` — https://dictation.assemblyai.com/v1/transcribe/live (docs: https://www.assemblyai.com/blog/dictation-api)
  - Pricing source: https://www.assemblyai.com/pricing
Notable capabilities:
  - Promptable speech language model: Universal-3 Pro (Feb 2026) introduced plain-language prompts controlling transcription (disfluencies, multilingual handling, PII, formatting); 3.5 Pro focuses on entities, rare words and domain terms with an LLM-based decoder. (https://www.assemblyai.com/blog/introducing-universal-3-pro)
  - Dictation API (polished text from short utterances) (found after launch): Launched 2026-09-15: up to 5 s audio per request (chunked upload), removes fillers, resolves self-corrections and fixes name spellings via `llm_instruction`, `keyterms_prompt` and `stt_prompt`; 0.36 s average response, 3.87% WER on short-form audio (vendor-cited), 19 languages, $0.62/hour all-in. Open-source MIT macOS demo app 'Blurt'. (https://www.assemblyai.com/blog/dictation-api)
  - Medical Mode (found after launch): `domain: medical-v1` for EN/ES/DE/FR clinical vocabulary; replaces deprecated Slam-1. (https://www.assemblyai.com/llms/models.md)


## bilibili (Index Team)

### IndexTTS-2 / IndexTTS-2.5 (bilibili)
**IndexTTS-2 / IndexTTS-2.5 (bilibili)** (bilibili (Index Team); current; audio/speech; released 2025-09-08; open weights) | Hugging Face: `IndexTeam/IndexTTS-2` | GitHub: https://github.com/index-tts/index-tts — The IndexTTS2 paper (arXiv June 2025) presents duration control as novel for AR TTS; 'first' not independently verified, so not flagged. Weights released 2025-09-08. Commercial use: contact indexspeech@bilibili.com.
  - Hugging Face: `IndexTeam/IndexTTS-2` — https://huggingface.co/IndexTeam/IndexTTS-2
  - GitHub — https://github.com/index-tts/index-tts
Notable capabilities:
  - Precise duration control in an autoregressive TTS: Lets users specify the exact number of speech tokens (useful for dubbing/lip-sync) while keeping AR naturalness, and disentangles speaker timbre from emotion (emotion from a separate reference audio or text). (https://huggingface.co/IndexTeam/IndexTTS-2)
  - IndexTTS-2.5 multilingual (found after launch): 2026-08-10 release adds Japanese, Spanish and Arabic to Chinese/English; speed 0.5-2x, Pinyin/CMU/Kana pronunciation control, RTF ~0.2 on RTX 4090. (https://github.com/index-tts/index-tts)


## Black Forest Labs

### FLUX.2 [klein] (4B / 9B)
**FLUX.2 [klein] (4B / 9B)** (Black Forest Labs; current; image-gen; released 2026-01-14; open weights) | BFL API: `flux-2-klein-4b`; BFL API: `flux-2-klein-9b` | Hugging Face: https://huggingface.co/black-forest-labs/FLUX.2-klein-4B; Hugging Face: https://huggingface.co/black-forest-labs/FLUX.2-klein-9B — Snapshots flux-2-klein-9b (fixed) and flux-2-klein-9b-preview (latest, KV caching). HF also hosts -base, fp8 and nvfp4 variants. Release date = HF repo creation date.
  - BFL API: `flux-2-klein-4b` — https://api.bfl.ai/v1/flux-2-klein-4b (docs: https://docs.bfl.ai/flux_2/flux2_overview)
  - BFL API: `flux-2-klein-9b` — https://api.bfl.ai/v1/flux-2-klein-9b (docs: https://docs.bfl.ai/flux_2/flux2_overview)
  - Hugging Face — https://huggingface.co/black-forest-labs/FLUX.2-klein-4B
  - Hugging Face — https://huggingface.co/black-forest-labs/FLUX.2-klein-9B
  - Pricing source: https://docs.bfl.ai/quick_start/pricing
Notable capabilities:
  - Sub-second generation and editing: Size-distilled FLUX.2 variants aimed at sub-second inference for both text-to-image and editing. (https://docs.bfl.ai/flux_2/flux2_overview)
  - Apache-2.0 open weights (4B): 4B checkpoint is Apache 2.0 - commercially usable open weights; base (undistilled) checkpoints published for fine-tuning/LoRA training. (https://huggingface.co/black-forest-labs/FLUX.2-klein-4B)
  - KV-cached 9B variant (found after launch): flux-2-klein-9b-preview / FLUX.2-klein-9b-kv (Mar 2026) add KV caching for faster multi-reference editing. (https://docs.bfl.ai/flux_2/flux2_overview)

### FLUX.2 [max]
**FLUX.2 [max]** (Black Forest Labs; current; image-gen; released 2025-12) | BFL API: `flux-2-max` | Web app: https://playground.bfl.ai — Release month (Dec 2025) not confirmed on an official page. Endpoint confirmed in https://api.bfl.ai/openapi.json.
  - BFL API: `flux-2-max` — https://api.bfl.ai/v1/flux-2-max (docs: https://docs.bfl.ai/flux_2/flux2_overview)
  - Web app — https://playground.bfl.ai
  - Pricing source: https://docs.bfl.ai/quick_start/pricing
Notable capabilities:
  - Grounded generation with web search: Can pull real-time web context (grounding search) into generations, e.g. current events or real products. (https://bfl.ai/models/flux-2-max)
  - Highest editing consistency in FLUX.2: Top FLUX.2 tier for prompt following, style fidelity, character consistency and retexturing/product photography. (https://bfl.ai/models/flux-2-max)

### FLUX.2 [dev]
**FLUX.2 [dev]** (Black Forest Labs; current; image-gen; released 2025-11-25; open weights) | Hugging Face: https://huggingface.co/black-forest-labs/FLUX.2-dev; Hugging Face (NVFP4): https://huggingface.co/black-forest-labs/FLUX.2-dev-NVFP4 — Open weights only (no /v1/flux-2-dev endpoint in BFL API openapi.json); commercial use needs a BFL license (https://bfl.ai/licensing). Hosted by many third parties. Pricing n/a.
  - Hugging Face — https://huggingface.co/black-forest-labs/FLUX.2-dev
  - Hugging Face (NVFP4) — https://huggingface.co/black-forest-labs/FLUX.2-dev-NVFP4
Notable capabilities:
  - 32B open-weight generation + multi-reference editing: 32B open-weight model doing text-to-image, single- and multi-reference editing in one checkpoint; BFL claims it beats all open-weight alternatives. (https://bfl.ai/blog/flux-2)
  - VLM-conditioned rectified flow transformer: Pairs a Mistral-3 24B vision-language model with a rectified flow transformer for world knowledge and prompt understanding. (https://bfl.ai/blog/flux-2)

### FLUX.2 [pro]
**FLUX.2 [pro]** (Black Forest Labs; current; image-gen; released 2025-11-25) | BFL API: `flux-2-pro` | Web app: https://playground.bfl.ai — flux-2-pro is a fixed snapshot; flux-2-pro-preview tracks the latest [pro]. Siblings: flux-2-flex (from $0.05, step/guidance control), flux-2-max. Uses Mistral-3 24B VLM + rectified flow transformer.
  - BFL API: `flux-2-pro` — https://api.bfl.ai/v1/flux-2-pro (docs: https://docs.bfl.ai/flux_2/flux2_overview)
  - Web app — https://playground.bfl.ai
  - Pricing source: https://docs.bfl.ai/quick_start/pricing
Notable capabilities:
  - Multi-reference editing (up to 10 images): Generates and edits with up to 10 reference images for character/product/style consistency, in one model with text-to-image. (https://bfl.ai/blog/flux-2)
  - 4MP editing and production-grade typography: Image editing up to 4 megapixels; reliable fine text for infographics, memes and UI mockups. (https://bfl.ai/blog/flux-2)

### FLUX 3
**FLUX 3** (Black Forest Labs; preview; video-gen; released 2026-07-23) | BFL API: `flux-3-video` | Hugging Face (FLUX 3 Action open weights): https://huggingface.co/black-forest-labs/flux-3-action-base — Early access at launch (2026-07-23); FLUX 3 Image announced 'in coming weeks' and open FLUX 3 [dev] planned later in 2026 - not verified as released. Action weights: flux-3-action-base/-so101/-droid (HF, 2026-09-22).
  - BFL API: `flux-3-video` — https://api.bfl.ai/v1/flux-3-video (docs: https://docs.bfl.ai/flux_3/flux3_overview)
  - Hugging Face (FLUX 3 Action open weights) — https://huggingface.co/black-forest-labs/flux-3-action-base
  - Pricing source: https://docs.bfl.ai/quick_start/pricing
Notable capabilities:
  - Unified image/video/audio/action model: Single architecture jointly trained on images, video, audio and robot action prediction; each modality said to strengthen the others. (https://www.globenewswire.com/news-release/2026/07/23/3332364/0/en/black-forest-labs-unveils-flux-3-a-new-multimodal-frontier-model-for-visual-intelligence.html)
  - Video with native synced audio: Text/image-to-video up to ~20 s with optional in-sync audio, plus video continuation and video editing (/v1/flux-tools/video-edit-v1). (https://docs.bfl.ai/flux_3/flux3_overview)
  - FLUX 3 Action for robotics: Video-prediction engine reused for robot control (FLUX-mimic with mimic robotics, tested by Audi); open-weight Action checkpoints on HF (FLUX Kommunity license). (https://docs.bfl.ai/flux_3/flux3_action_overview)

### FLUX.1 Kontext [pro] / [max]
**FLUX.1 Kontext [pro] / [max]** (Black Forest Labs; legacy; image-gen; released 2025-05-29) | BFL API: `flux-kontext-pro`; BFL API: `flux-kontext-max` | Hugging Face (open Kontext [dev]): https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev — Previous generation (BFL pricing page lists FLUX.1 as 'previous generation'); still served. Release date from memory of BFL launch (May 2025), not re-verified today. Also still served: flux-pro-1.1 ($0.04), flux-pro-1.1-ultra ($0.06).
  - BFL API: `flux-kontext-pro` — https://api.bfl.ai/v1/flux-kontext-pro (docs: https://docs.bfl.ai/kontext/kontext_overview)
  - BFL API: `flux-kontext-max` — https://api.bfl.ai/v1/flux-kontext-max (docs: https://docs.bfl.ai/kontext/kontext_overview)
  - Hugging Face (open Kontext [dev]) — https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev
  - Pricing source: https://docs.bfl.ai/quick_start/pricing
Notable capabilities:
  - In-context image editing: Text-instructed edits of an input image with character consistency across iterative edits; one model for generation and editing. (https://docs.bfl.ai/kontext/kontext_overview)
  - Open-weight editing sibling: FLUX.1 Kontext [dev] released as open weights (non-commercial) for local instruction-based editing. (https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev)


## Boson AI

### Boson AI Higgs Audio v3 (Higgs TTS 3 4B / Higgs STT 3)
**Boson AI Higgs Audio v3 (Higgs TTS 3 4B / Higgs STT 3)** (Boson AI; current; audio/speech; released 2026-06-04; open weights) | Hugging Face: `bosonai/higgs-audio-v3-tts-4b` | GitHub: https://github.com/boson-ai/higgs-audio — TTS weights non-commercial; production/hosted use needs a Boson commercial license or the Boson API (pricing not found). Also mirrored as bosonai/higgs-tts-3-4b. Predecessor Higgs Audio v2 (2025, Apache-2.0-style) on the same GitHub.
  - Hugging Face: `bosonai/higgs-audio-v3-tts-4b` — https://huggingface.co/bosonai/higgs-audio-v3-tts-4b
  - GitHub — https://github.com/boson-ai/higgs-audio
  - SGLang-Omni (docs: https://www.lmsys.org/blog/2026-06-04-higgs-audio-v3-tts/)
Notable capabilities:
  - 102-language expressive TTS with zero-shot cloning: ~4B AR decoder (24 kHz, 8 codebooks); 85 languages at production quality (WER/CER <5%), 17 usable; inline control of emotion, style, prosody, pauses and sound effects; 8K-token context; sub-second TTFA streaming. (https://huggingface.co/bosonai/higgs-audio-v3-tts-4b)
  - Higgs STT 3 (API): Speech-to-text model (2026-03-18) for 94 languages; 1.55% WER on LibriSpeech test-clean vs 2.10% for Whisper-large-v3 (company figures). No open weights found. (https://www.boson.ai/blog/higgs-audio-v3-stt)


## Boston Dynamics

### Atlas Large Behavior Model (Boston Dynamics + TRI LBM)
**Atlas Large Behavior Model (Boston Dynamics + TRI LBM)** (Boston Dynamics; current; robotics; released 2025-08-20) | Not available (internal research policy for Atlas): https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/; TRI LBM Eval (open simulation benchmark, not the model): https://github.com/ToyotaResearchInstitute/lbm_eval — Research collaboration announced Aug 2025 (Toyota release: https://newsroom.toyota.eu/ai-powered-robot-by-boston-dynamics-and-toyota-research-institute-takes-a-key-step-towards-general-purpose-humanoids/). The production electric Atlas (CES 2026) also integrates Google DeepMind foundation models (Gemini Robotics); Hyundai trains Atlas on parts logistics at its Georgia RMAC (2026-09-22). No public weights or API for the Atlas LBM. Exact announcement day (2025-08-20) is from press coverage dated 2025-08-20/21.
  - Not available (internal research policy for Atlas) — https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/
  - TRI LBM Eval (open simulation benchmark, not the model) — https://github.com/ToyotaResearchInstitute/lbm_eval
Notable capabilities:
  - One language-conditioned policy for whole-body loco-manipulation: A single end-to-end policy maps images, proprioception and language to actions for the full 50-DoF Atlas at 30 Hz, combining stepping, crouching and center-of-mass shifts with dexterous manipulation in long-horizon tasks, replacing separate walking and manipulation controllers. (https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/)
  - Diffusion Transformer with flow matching: 450M-parameter Diffusion Transformer trained with a flow-matching objective, predicting 48-step action chunks (1.6 s); trained on Atlas teleop data, the Atlas Manipulation Test Stand, TRI's Ramen dataset and simulation co-training. (https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/)
  - Inference-time speed-up: Policies can run 1.5-2x faster than the human demos at inference time without retraining by rescaling action timing. (https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/)
  - Pretraining cuts task data by up to 80%: TRI's LBM study (~1,700 h of robot data, 1,800 real and 47,000+ sim rollouts) found pretrained LBMs need up to 80% less task-specific data. (https://toyotaresearchinstitute.github.io/lbm1/)


## BreezeBlue

### Breeze TTS 2
**Breeze TTS 2** (BreezeBlue; current; audio/speech; released 2026-08-25; open weights) | Hugging Face: `BreezeBlue/Breeze-TTS-2` | GitHub: https://github.com/breezeblue-ai/breeze-tts; BreezeBlue (hosted / commercial license): https://breezeblue.ai — Model card lists English + Chinese; the Artificial Analysis post mentions 50 languages (possibly the hosted model) — unresolved. Needs 12 GB VRAM (24 GB recommended), CUDA/Linux. Weights are NOT commercially usable without a BreezeBlue subscription. Some secondary blogs claim it is the 'first open-weight model to beat ElevenLabs' flagship' — unverified and contradicted by the AA leaderboard (Eleven v4 far ahead).
  - Hugging Face: `BreezeBlue/Breeze-TTS-2` — https://huggingface.co/BreezeBlue/Breeze-TTS-2
  - GitHub — https://github.com/breezeblue-ai/breeze-tts
  - BreezeBlue (hosted / commercial license) — https://breezeblue.ai
Notable capabilities:
  - #1 open-weights TTS on Artificial Analysis (found after launch): ~1,206-1,215 Elo in the Artificial Analysis Speech Arena, ~90 points above Fish Audio S2 Pro, #6 overall at launch — the leading open-weights TTS as of Sept 2026. (https://x.com/ArtificialAnlys/status/2092399623839326550)
  - Clone + design + direct in one 3B checkpoint, <40 ms TTFA: Voice cloning from reference audio, voice design from text descriptions, voice direction (tone/emotion keeping identity), vocal events (laughs, coughs); streaming TTFA under 40 ms on H100 with fast path, RTF 0.32. (https://huggingface.co/BreezeBlue/Breeze-TTS-2)


## ByteDance

### SeedRealtime (Doubao realtime audio-visual model)
**SeedRealtime (Doubao realtime audio-visual model)** (ByteDance; current; audio/speech; released 2026-08-05) | Web app (Doubao / Dola): https://dola.com/chat; BytePlus Playground: https://ai.byteplus.com/en/playground — Deployed at scale in the Doubao app (Dola internationally). No public API model id, pricing or benchmark numbers published; Volcengine offers a separate Doubao end-to-end realtime dialogue API (/api/v3/realtime/dialogue) whose relation to SeedRealtime is unverified. Some press calls it the first model to watch, listen and speak simultaneously; not claimed by ByteDance, and Gemini Live / GPT-Realtime already accepted video.
  - Web app (Doubao / Dola) — https://dola.com/chat
  - BytePlus Playground — https://ai.byteplus.com/en/playground
Notable capabilities:
  - Native audio-visual full-duplex LLM: Single end-to-end model perceives continuous audio, video and text streams while listening and speaking (no ASR/VLM/TTS cascade); resolves homophones from visual context and temporal references to what it sees. (https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction)
  - Proactive turn-taking: ByteDance says it halves audio-visual conversational pacing problems vs cascaded systems (fewer cut-offs, slow replies, false triggers) and can speak up proactively. (https://seed.bytedance.com/en/SeedRealtime)

### Seed Audio 1.0
**Seed Audio 1.0** (ByteDance; current; audio/speech; released 2026-07-20) | BytePlus (Seed Speech console): https://console.byteplus.com/voice/new/setting/activate?projectName=default — API model id and pricing not found. Related ByteDance speech stack: Seed-TTS 2.0 / Doubao TTS 2.0 (Oct 2025), Doubao-Seed-ASR-2.0, Seed LiveInterpret 2.0 (2025-07-24, zh<->en simultaneous interpretation with voice cloning, ~2.5-3 s lag; see entry 2025-07-24-bytedance-seed-liveinterpret-2) on Volcengine/BytePlus. Comparable: Qwen-Audio-3.1-TTS-Next, StepAudio 3 Gen.
  - BytePlus (Seed Speech console) — https://console.byteplus.com/voice/new/setting/activate?projectName=default
Notable capabilities:
  - Unified speech + SFX + ambience generation: Jointly models voice, sound effects and ambience in one framework for film-grade audio; multi-character dialogue with prompt-level timing control at 100 ms precision; up to 2 min per generation with continuation. (https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model)
  - 20+ languages: Including zh, en, ja, ko, es, id, de, fr, th, vi; most languages MOS > 4.0 in ByteDance's evaluation. (https://seed.bytedance.com/en/seedaudio1_0)


## Canopy Labs

### Canopy Labs Orpheus TTS (3B)
**Canopy Labs Orpheus TTS (3B)** (Canopy Labs; current; audio/speech; released 2025-03; open weights) | Hugging Face: `canopylabs/orpheus-tts-0.1-finetune-prod`; Hugging Face: `canopylabs/orpheus-3b-0.1-ft`; Groq: `canopylabs/orpheus-v1-english`; Groq: `canopylabs/orpheus-arabic-saudi` | Together AI: https://www.together.ai/models/orpheus-tts — 8 English preset voices (tara, leah, jess, leo, dan, mia, zac, zoe); multilingual research release (7 language pairs) April 2025. Groq deployed two variants on 2026-01-13 (press: $22 per 1M characters, not verified on Groq pricing page).
  - Hugging Face: `canopylabs/orpheus-tts-0.1-finetune-prod` — https://github.com/canopyai/Orpheus-TTS
  - Hugging Face: `canopylabs/orpheus-3b-0.1-ft` — https://huggingface.co/canopylabs/orpheus-3b-0.1-ft
  - Groq: `canopylabs/orpheus-v1-english` — https://api.groq.com/openai/v1/audio/speech (docs: https://console.groq.com/docs/text-to-speech)
  - Groq: `canopylabs/orpheus-arabic-saudi` — https://api.groq.com/openai/v1/audio/speech
  - Together AI — https://www.together.ai/models/orpheus-tts
Notable capabilities:
  - LLM-backbone TTS with emotion tags: Llama-3B-based speech LLM trained on 100k+ h English; tags <laugh>, <chuckle>, <sigh>, <cough>, <sniffle>, <groan>, <yawn>, <gasp>; ~200 ms streaming latency (~100 ms with input streaming); zero-shot cloning via pretrained model. (https://github.com/canopyai/Orpheus-TTS)


## Cartesia

### Cartesia Sonic-3.6
**Cartesia Sonic-3.6** (Cartesia; current; audio/speech; released 2026-08-27) | Cartesia API: `sonic-3.6`; Cartesia API (pinned snapshot): `sonic-3.6-2026-08-27` | Web app: https://play.cartesia.ai — Beta 2026-08-17, GA snapshot 2026-08-27. Header `Cartesia-Version: 2026-08-14`. Fully backwards compatible with Sonic-3.5 (snapshot 2026-05-04, which led AA's Controlled Voice Arena at its 2026-07-08 launch with 1,122 Elo). Scored 0.840 (#5) on Hume's Real-World VoiceEQ leaderboard (2026-09-24). sonic-3 snapshots (2025-10-27, 2026-01-12), sonic-2 and sonic-turbo sunset 2026-10-20. `sonic-preview` = beta channel; `sonic-latest` alias deprecated. Exact per-character USD price is plan-dependent (credits); figure above is derived. Also on AWS SageMaker JumpStart (Sonic 3, Feb 2026).
  - Cartesia API: `sonic-3.6` — https://api.cartesia.ai/tts/bytes (also /tts/sse and WebSocket) (docs: https://docs.cartesia.ai/build-with-cartesia/tts-models/latest)
  - Cartesia API (pinned snapshot): `sonic-3.6-2026-08-27` — https://api.cartesia.ai/tts/bytes
  - Web app — https://play.cartesia.ai
  - Pricing source: https://cartesia.ai/pricing
Notable capabilities:
  - State-space-model TTS, sub-90 ms: Built on state space models (SSMs); replies in under 90 ms and generates ~132 chars/s (nearly 2x Sonic 3 Conversational). Listeners preferred it over Sonic-3.5 in up to 93% of blind tests across 15 locales. (https://www.cartesia.ai/blog/sonic-3.6)
  - 44 languages with instant cloning: Adds Odia and Urdu to Sonic-3.5's 42 languages; instant voice cloning; locale-aware reading of dates/numbers; confirmation codes and heteronyms without preprocessing. (https://docs.cartesia.ai/build-with-cartesia/tts-models/latest)
  - Multilingual Voices (one voice, 25 languages) (found after launch): Launched 2026-09-23 on Sonic-3.6: 50+ library voices each speak up to 25 languages natively, and custom clones from ~10 s of audio carry their identity across languages via a `locale` parameter; native speakers rate each variant for accent and localization of dates, numbers and currency. (https://www.cartesia.ai/blog/multilingual-voices)
  - Top-2 on Artificial Analysis Speech Arena (found after launch): Ranked #1 (Elo ~1279) on the Artificial Analysis TTS leaderboard in mid/late Sept 2026, then #2 (Elo 1275) behind Eleven v4 after 2026-09-28. (https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice)

### Cartesia Ink-2 (streaming STT)
**Cartesia Ink-2 (streaming STT)** (Cartesia; current; audio/speech; released 2026-07-09) | Cartesia API: `ink-2`; Cartesia API (beta): `ink-preview` — Launched English-only (blog 2026-07-09); the current stable `ink-2` snapshot is dated 2026-09-17 and supports English, French, Hindi, Japanese, Spanish. Some press dates an earlier Ink 2 release to May 2026 (unverified). Query params: model, encoding, sample_rate, cartesia_version=2026-08-14; send `finalize` when user stops. Older model: ink-whisper (1 credit/s streaming). Ink-2 credit price not found on pricing page (plans list included STT hours).
  - Cartesia API: `ink-2` — wss://api.cartesia.ai/stt/websocket?model=ink-2 (docs: https://docs.cartesia.ai/build-with-cartesia/stt/latest)
  - Cartesia API (beta): `ink-preview` — wss://api.cartesia.ai/stt/websocket
Notable capabilities:
  - Built-in semantic turn detection: Emits turn.start / turn.update / turn.eager_end / turn.resume / turn.end events so agents need no separate VAD; 89% precision, 93% F1 on endpointing; ~0.1 s time-to-final-transcript. (https://www.cartesia.ai/blog/introducing-ink-2)
  - #1 streaming WER on Artificial Analysis at launch: 3.4% WER on AA-AgentTalk, ranked #1 on Artificial Analysis's streaming STT leaderboard (company claim, July 2026). (https://www.cartesia.ai/blog/introducing-ink-2)
  - Keyterm prompting (found after launch): Keyterm prompting and configurable turn detection added 2026-08-11. (https://www.cartesia.ai/blog)


## Cohere

### Command A+
**Command A+** (Cohere; current; reasoning-llm; released 2026-05-20; open weights) | ctx 128,000 | $0.3 in / $1.5 out per 1M tokens (USD) on OpenRouter; Cohere first-party price not verified | Cohere API: `command-a-plus-05-2026`; OpenRouter: `cohere/command-a-plus` | Hugging Face: https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4 — Also HF CohereLabs/command-a-plus-05-2026-bf16 and -fp8. OpenRouter lists 192K context vs 128K in Cohere docs. Cohere pricing page did not list per-token price.
  - Cohere API: `command-a-plus-05-2026` — https://api.cohere.com/v2/chat (docs: https://docs.cohere.com/docs/command-a-plus)
  - OpenRouter: `cohere/command-a-plus` — https://openrouter.ai/cohere/command-a-plus
  - Hugging Face — https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4
  - Pricing source: https://openrouter.ai/cohere/command-a-plus
Notable capabilities:
  - Cohere's first MoE model: 218B total / 25B active mixture-of-experts combining vision, agentic and reasoning capabilities in one model. (https://docs.cohere.com/docs/models)
  - Apache 2.0 enterprise model on 1 B200: Open weights under Apache 2.0 (earlier Command A was CC-BY-NC); W4A4 build runs on 1x B200 or 2x H100. (https://docs.cohere.com/docs/command-a-plus)
  - 48 languages: Supports 48 languages including all official EU languages, with configurable reasoning. (https://docs.cohere.com/docs/command-a-plus)

### Cohere Rerank 4 (Pro / Fast)
**Cohere Rerank 4 (Pro / Fast)** (Cohere; current; embedding; released 2025-12-11) | ctx 32,000 | Cohere API: `rerank-v4.0-pro`; Cohere API (fast): `rerank-v4.0-fast`; OpenRouter: `cohere/rerank-4-pro` — Reranker (scores query-document relevance). Previous: rerank-v3.5 (Bedrock cohere.rerank-v3-5:0). Release date from third-party listing; pricing not verified (OpenRouter ~$0.0025/search reported, not checked).
  - Cohere API: `rerank-v4.0-pro` — https://api.cohere.com/v2/rerank (docs: https://docs.cohere.com/docs/models)
  - Cohere API (fast): `rerank-v4.0-fast` — https://api.cohere.com/v2/rerank
  - OpenRouter: `cohere/rerank-4-pro` — https://openrouter.ai/cohere/rerank-4-pro
Notable capabilities:
  - 32K-context reranking: Rerank window grew from 4K (v3.5) to 32K tokens, so whole long documents can be scored. (https://docs.cohere.com/docs/models)
  - Pro / Fast tiers: Two variants: pro for best accuracy, fast for latency-sensitive search. (https://docs.cohere.com/docs/models)

### Cohere Embed v4
**Cohere Embed v4** (Cohere; current; embedding; released 2025-04) | ctx 128,000 | Cohere API: `embed-v4.0`; AWS Bedrock: `cohere.embed-v4:0` — Output is vectors (modality_out text used as placeholder). Release month (Apr 2025) from memory, not re-verified. Pricing not verified (Cohere pricing page shows only Model Vault hourly rates).
  - Cohere API: `embed-v4.0` — https://api.cohere.com/v2/embed (docs: https://docs.cohere.com/docs/cohere-embed)
  - AWS Bedrock: `cohere.embed-v4:0` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-cohere-embed-v4.html)
Notable capabilities:
  - Interleaved text+image (PDF) embeddings: Embeds text, images and mixed text/image documents such as PDFs into one vector space. (https://docs.cohere.com/docs/cohere-embed)
  - 128K-token input with Matryoshka dims: Up to 128K tokens per input; output dimension selectable 256/512/1024/1536. (https://docs.cohere.com/docs/models)

### Command A (03-2025) and variants
**Command A (03-2025) and variants** (Cohere; legacy; llm; released 2025-03; open weights) | ctx 256,000 | $2.5 in / $10 out per 1M tokens (USD) on OpenRouter; Cohere first-party price not verified | Cohere API: `command-a-03-2025`; OpenRouter: `cohere/command-a` | Hugging Face: https://huggingface.co/CohereLabs/c4ai-command-a-03-2025 — Superseded by Command A+ (May 2026). Variants listed in notes/capabilities share this file. Weights are non-commercial (CC-BY-NC).
  - Cohere API: `command-a-03-2025` — https://api.cohere.com/v2/chat (docs: https://docs.cohere.com/docs/command-a)
  - OpenRouter: `cohere/command-a` — https://openrouter.ai/cohere/command-a
  - Hugging Face — https://huggingface.co/CohereLabs/c4ai-command-a-03-2025
  - Pricing source: https://openrouter.ai/cohere/command-a
Notable capabilities:
  - Enterprise model on two GPUs: 111B model that runs on only two A100/H100 GPUs, 150% higher throughput than Command R+ 08-2024. (https://docs.cohere.com/docs/command-a)
  - Specialized variants (found after launch): Separate ids command-a-reasoning-08-2025 (256K/32K out), command-a-vision-07-2025 (image input) and command-a-translate-08-2025 (23-language MT). (https://docs.cohere.com/docs/models)


## Deepgram

### Deepgram Flux TTS
**Deepgram Flux TTS** (Deepgram; current; audio/speech; released 2026-08-12) | Deepgram API (real-time): `flux-haley-en`; Deepgram API (batch): `flux-{voice}-en` — English only (39 voices; American, British, Irish, Australian, Indian, Singaporean, Filipino accents); use Aura-2 for other languages. Self-hosted GA 2026-08-26; speed 0.5-1.5 and expressivity -2..2 controls. Launched alongside Deepgram passing $100M ARR.
  - Deepgram API (real-time): `flux-haley-en` — wss://api.deepgram.com/v2/speak?model=flux-haley-en (docs: https://developers.deepgram.com/docs/flux-tts/overview)
  - Deepgram API (batch): `flux-{voice}-en` — https://api.deepgram.com/v2/speak (docs: https://developers.deepgram.com/docs/flux-tts/voices)
  - Pricing source: https://deepgram.com/pricing
Notable capabilities:
  - Conversation-native TTS: Keeps context and voice consistency across turns of a conversation instead of treating each sentence in isolation; turn lifecycle events; on Interrupt reports exactly what the user heard (`text_spoken`). (https://deepgram.com/learn/text-to-speech-comes-of-age-deepgram-launches-conversation-native-speech)
  - ~80 ms response, structured-content accuracy: Starts responding in as little as 80 ms under production load; tuned for account numbers, alphanumerics, drug names and money amounts. (https://deepgram.com/learn/text-to-speech-comes-of-age-deepgram-launches-conversation-native-speech)

### Deepgram Flux (conversational STT, English + Multilingual)
**Deepgram Flux (conversational STT, English + Multilingual)** (Deepgram; current; audio/speech; released 2025-10-02) | Deepgram API: `flux-general-en`; Deepgram API: `flux-general-multi` — 'First' claims are Deepgram's own marketing (launched at VapiCon 2025-10-02 as 'world's first conversational speech recognition model'). Uses /v2/listen (not /v1). Mid-stream numeral toggle added 2026-09-25. Companion TTS: deepgram-flux-tts.
  - Deepgram API: `flux-general-en` — wss://api.deepgram.com/v2/listen (docs: https://developers.deepgram.com/docs/flux/quickstart)
  - Deepgram API: `flux-general-multi` — wss://api.deepgram.com/v2/listen (docs: https://developers.deepgram.com/docs/models-languages-overview)
  - Pricing source: https://deepgram.com/pricing
Notable capabilities:
  - [FIRST] Conversational speech recognition with model-native turn-taking: Recognition model itself decides end-of-turn using acoustic + semantic cues (~260 ms end-of-turn detection), with EagerEndOfTurn events to start the LLM early; tunable eot_threshold, eager_eot_threshold, eot_timeout_ms. (https://deepgram.com/learn/introducing-flux-conversational-speech-recognition)
  - [FIRST] Multilingual conversational STT with in-call code-switching (found after launch): Flux Multilingual (GA 2026-04-29): English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch with automatic language switching mid-conversation; turn detection under 400 ms. Billed by Deepgram as the world's first multilingual conversational speech recognition model. (https://deepgram.com/learn/deepgram-launches-flux-multilingual-press-release)

### Deepgram Aura-2
**Deepgram Aura-2** (Deepgram; current; audio/speech; released 2025-04-15) | Deepgram API: `aura-2-thalia-en` — For English voice agents Deepgram now recommends Flux TTS (deepgram-flux-tts); Aura-2 remains the multilingual option.
  - Deepgram API: `aura-2-thalia-en` — https://api.deepgram.com/v1/speak?model=aura-2-thalia-en (docs: https://developers.deepgram.com/docs/tts-models)
  - Pricing source: https://deepgram.com/pricing
Notable capabilities:
  - Enterprise TTS with deployable runtime: Sub-200 ms TTFB, cloud/VPC/on-prem deployment; model id pattern aura-2-{voice}-{lang}. (https://deepgram.com/learn/introducing-aura-2-enterprise-text-to-speech)
  - 7 languages, EN/ES code-switching voices (found after launch): English, Spanish, German, French, Dutch, Italian, Japanese; several Spanish voices code-switch with English. (https://developers.deepgram.com/docs/tts-models)

### Deepgram Nova-3 (incl. Medical / Pharma)
**Deepgram Nova-3 (incl. Medical / Pharma)** (Deepgram; current; audio/speech; released 2025-02-12) | Deepgram API: `nova-3`; Deepgram API: `nova-3-medical`; Deepgram API: `nova-3-pharma` — Release date 2025-02-12 from Deepgram's Nova-3 launch (not re-checked today). Previous gen nova-2 and variants still served. Deepgram also hosts Whisper (whisper-large, $0.0048/min).
  - Deepgram API: `nova-3` — https://api.deepgram.com/v1/listen (REST) / wss://api.deepgram.com/v1/listen (docs: https://developers.deepgram.com/docs/models-languages-overview)
  - Deepgram API: `nova-3-medical`
  - Deepgram API: `nova-3-pharma`
  - Pricing source: https://deepgram.com/pricing
Notable capabilities:
  - Keyterm prompting, 90+ languages (found after launch): Nova-3 general supports 90+ languages incl. multilingual code-switching mode; languages added continuously through 2026 (e.g. Kazakh 2026-09-03, Assamese/Mongolian/Pashto 2026-08-27). (https://developers.deepgram.com/changelog)
  - Domain variants (found after launch): nova-3-medical (upgraded batch model May 2026) and nova-3-pharma (English pharmaceutical model, 2026-09-17). (https://developers.deepgram.com/changelog)


## DeepSeek

### DeepSeek-V4.1-Flash
**DeepSeek-V4.1-Flash** (DeepSeek; current; reasoning-llm; released 2026-09-10; open weights) | ctx 1,000,000 | $0.3 in / $1.2 out per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.15, output 0.6, cache hit 0.003). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri | DeepSeek API: `deepseek-flash`; DeepSeek API (Anthropic format): `deepseek-flash`; Alibaba Cloud Model Studio: `deepseek-v4.1-flash`; OpenRouter: `deepseek/deepseek-v4.1-flash` | Hugging Face: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash; Web app: https://chat.deepseek.com — Call as deepseek-flash. Legacy ids deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at Flash price. Knowledge cutoff not published.
  - DeepSeek API: `deepseek-flash` — https://api.deepseek.com (docs: https://api-docs.deepseek.com/quick_start/pricing)
  - DeepSeek API (Anthropic format): `deepseek-flash` — https://api.deepseek.com/anthropic (docs: https://api-docs.deepseek.com/quick_start/pricing)
  - Alibaba Cloud Model Studio: `deepseek-v4.1-flash` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - OpenRouter: `deepseek/deepseek-v4.1-flash` — https://openrouter.ai/deepseek/deepseek-v4.1-flash
  - Hugging Face — https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
  - Web app — https://chat.deepseek.com
  - Pricing source: https://api-docs.deepseek.com/quick_start/pricing
Notable capabilities:
  - Native vision in the Flash tier: First DeepSeek Flash model with native multimodal (image) understanding built in; replaced the separate V4-Flash-Vision-Exp. (https://api-docs.deepseek.com/updates)
  - Causal Encoder-Decoder (CED) architecture: 552B-backbone MoE that activates only ~8B params per token in prefill and ~16B in decode, aimed at input-heavy agentic workloads. (https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
  - Tiny KV cache (CSA2 + FP4 KV): Compressed Sparse Attention 2 and FP4 main KV cache cut the global KV cache to ~890 bytes/token, about 1/4 of V4-Flash. (https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
  - Hybrid thinking with effort levels: One model id serves thinking (default) and non-thinking modes; reasoning effort low/high/max. (https://api-docs.deepseek.com/updates)
  - Multiple API protocols: Same model served via OpenAI Chat Completions, OpenAI Responses (Codex-adapted) and Anthropic Messages formats. (https://api-docs.deepseek.com/quick_start/pricing)

### DeepSeek-V4-Pro
**DeepSeek-V4-Pro** (DeepSeek; current; reasoning-llm; released 2026-04-24; open weights) | ctx 1,000,000 | $1.32 in / $3.96 out per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.66, output 1.98, cache hit 0.022). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri | DeepSeek API: `deepseek-v4-pro`; DeepSeek API (Anthropic format): `deepseek-v4-pro`; Alibaba Cloud Model Studio: `deepseek-v4-pro-0813`; OpenRouter: `deepseek/deepseek-v4-pro-0813`; OpenRouter (preview 0423): `deepseek/deepseek-v4-pro` | Hugging Face: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813; Web app: https://chat.deepseek.com — Preview 2026-04-24, GA snapshot DeepSeek-V4-Pro-0813 on 2026-08-13 (same id deepseek-v4-pro). Text-only (no vision). DeepSeek said service continues past 2026-09-14 until further notice. Knowledge cutoff not published.
  - DeepSeek API: `deepseek-v4-pro` — https://api.deepseek.com (docs: https://api-docs.deepseek.com/quick_start/pricing)
  - DeepSeek API (Anthropic format): `deepseek-v4-pro` — https://api.deepseek.com/anthropic (docs: https://api-docs.deepseek.com/quick_start/pricing)
  - Alibaba Cloud Model Studio: `deepseek-v4-pro-0813` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - OpenRouter: `deepseek/deepseek-v4-pro-0813` — https://openrouter.ai/deepseek/deepseek-v4-pro-0813
  - OpenRouter (preview 0423): `deepseek/deepseek-v4-pro` — https://openrouter.ai/deepseek/deepseek-v4-pro
  - Hugging Face — https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813
  - Web app — https://chat.deepseek.com
  - Pricing source: https://api-docs.deepseek.com/quick_start/pricing
Notable capabilities:
  - Open-weight 1.6T MoE with 1M context: 1.6T total / 49B active parameters, MIT license, 1M-token context (paper: 'Towards Highly Efficient Million-Token Context Intelligence'). (https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
  - Agentic GA upgrade (0813) (found after launch): GA release greatly strengthened agent performance in production (e.g. Terminal Bench 2.1 87.9, Toolathlon-Verified 74.1 per DeepSeek). (https://api-docs.deepseek.com/updates)
  - Reasoning effort low/high/max (found after launch): Thinking mode supports three effort levels; non-thinking mode also available. (https://api-docs.deepseek.com/updates)
  - Native OpenAI Responses API + Codex (found after launch): DeepSeek API natively speaks the Responses API format and is adapted for Codex; Anthropic Messages format also supported. (https://api-docs.deepseek.com/updates)
  - DSpark speculative decoding module (found after launch): 0813 weights ship with an attached DSpark speculative-decoding module for faster inference. (https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813)

### DeepSeekMath-V2
**DeepSeekMath-V2** (DeepSeek; current; reasoning-llm; released 2025-11-27; open weights) | Hugging Face: https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 — 685B open-weights (Apache 2.0) math prover built on DeepSeek-V3.2-Exp-Base; inference uses the DeepSeek-V3.2-Exp code. No first-party API endpoint verified. 'first' = first open-weights model at IMO-gold level (per the paper's claims).
  - Hugging Face — https://huggingface.co/deepseek-ai/DeepSeek-Math-V2
Notable capabilities:
  - [FIRST] Self-verifiable proof generation: Generator trained against an LLM proof verifier and meta-verifier; reached IMO 2025 / CMO 2024 gold level and 118/120 on Putnam 2024 with scaled test-time compute. (https://arxiv.org/abs/2511.22570)

### DeepSeek-V3.2
**DeepSeek-V3.2** (DeepSeek; legacy; reasoning-llm; released 2025-12-01; open weights) | DeepSeek API (retired): `deepseek-chat / deepseek-reasoner (no longer serve V3.2)`; OpenRouter: `deepseek/deepseek-v3.2` | Hugging Face: https://huggingface.co/deepseek-ai/DeepSeek-V3.2; Hugging Face (Speciale): https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale — API aliases deepseek-chat/deepseek-reasoner moved to V4-Flash on 2026-04-24 and were scheduled for discontinuation on 2026-07-24; V3.2 now only via open weights/third parties. Pricing not verified (no first-party price).
  - DeepSeek API (retired): `deepseek-chat / deepseek-reasoner (no longer serve V3.2)` — https://api.deepseek.com (docs: https://api-docs.deepseek.com/updates)
  - OpenRouter: `deepseek/deepseek-v3.2` — https://openrouter.ai/deepseek/deepseek-v3.2
  - Hugging Face — https://huggingface.co/deepseek-ai/DeepSeek-V3.2
  - Hugging Face (Speciale) — https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale
Notable capabilities:
  - DeepSeek Sparse Attention (DSA): Introduced DSA (first in V3.2-Exp) to cut long-context attention compute while preserving quality. (https://huggingface.co/deepseek-ai/DeepSeek-V3.2)
  - Hybrid thinking/non-thinking in one model: deepseek-chat mapped to non-thinking mode and deepseek-reasoner to thinking mode of the same V3.2 weights. (https://api-docs.deepseek.com/updates)
  - V3.2-Speciale reasoning variant: Separate high-compute Speciale variant served briefly on a temporary endpoint (no tool calls) until 2025-12-15; weights released. (https://api-docs.deepseek.com/updates)


## Dyna Robotics

### DYNA-2 (World-Action Model)
**DYNA-2 (World-Action Model)** (Dyna Robotics; current; robotics; released 2026-08-10) | Dyna Robotics (commercial deployments): https://www.dyna.co/dyna-2 — Predecessor DYNA-1 (2025) runs in production in hotels, restaurants and laundromats (towel folding etc.). No API or weights; adapts to arms, humanoid prototypes and dexterous hands with hours of local fine-tuning. Figure (Helix 2.5), Generalist (GEN-1) and Dyna all reported human-video scaling in 2026, so 'first' claims overlap. Company-reported.
  - Dyna Robotics (commercial deployments) — https://www.dyna.co/dyna-2
Notable capabilities:
  - [FIRST] Human-to-robot scaling law: Pretrained on 1M+ hours of egocentric human video (~170 years); on-robot normalized score rose from 20% to 53% across 14 tasks as pretraining scaled from 1k to 1M hours. Dyna calls it the first scaling law demonstrated across the embodiment gap. (https://www.dyna.co/dyna-2)
  - World-action model: One video-diffusion (mixture-of-transformers, flow matching) model that denoises future video and an action chunk jointly or separately; one-step distilled video generation 90x faster than the teacher. (https://www.dyna.co/dyna-2)
  - Production quality gains: 87% zero-shot customer-quality pass rate at a customer deployment vs 46% for DYNA-1; 1.55x more task completions than DYNA-1; bottle-cap opening learned with 10 minutes of robot data. (https://www.prnewswire.com/news-releases/dyna-robotics-unveils-dyna-2-world-action-model-demonstrating-first-true-scaling-law-in-robotics-powered-entirely-by-human-data-302847114.html)


## ElevenLabs

### Eleven v4 / Eleven v4 Turbo
**Eleven v4 / Eleven v4 Turbo** (ElevenLabs; current; audio/speech; released 2026-09-28) | ElevenLabs API: `eleven_v4`; ElevenLabs API: `eleven_v4_turbo`; fal: `elevenlabs/tts/eleven-v4`; fal: `elevenlabs/tts/eleven-v4-turbo` | Web app (ElevenCreative): https://elevenlabs.io/app; Landing page / demos: https://elevenlabs.io/v4 — Launched 2026-09-28 (blog, YouTube 07:01 PT, X) in ElevenAgents, ElevenCreative and ElevenAPI, incl. free tier. eleven_v4: 10,000 chars/request; eleven_v4_turbo: no char limit listed on models page. Output formats MP3, WAV/PCM, u-law. Limitations: no Style/Speed sliders, no SSML (Stability + Similarity only); Voice Design voices may perform worse than with earlier models. Launch promo also: v4 free for Creator+ plans in ElevenCreative up to 2x monthly credits for two weeks. Third-party: on fal since launch day (fal X post https://x.com/fal/status/2104630460542325071): elevenlabs/tts/eleven-v4 at $0.08/1K chars and elevenlabs/tts/eleven-v4-turbo at $0.04/1K chars, the ElevenLabs list prices (fal model pages, checked 2026-09-29). Research led by Piotr Dabkowski (per press).
  - ElevenLabs API: `eleven_v4` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/overview/capabilities/text-to-speech/eleven-v4)
  - ElevenLabs API: `eleven_v4_turbo` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/models)
  - fal: `elevenlabs/tts/eleven-v4` (docs: https://fal.ai/models/elevenlabs/tts/eleven-v4)
  - fal: `elevenlabs/tts/eleven-v4-turbo` (docs: https://fal.ai/models/elevenlabs/tts/eleven-v4-turbo)
  - Web app (ElevenCreative) — https://elevenlabs.io/app
  - Landing page / demos — https://elevenlabs.io/v4
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Context-aware "performed" delivery (new architecture): Entirely new TTS architecture that 'reads a script the way a voice actor would', interpreting tone, pacing, emotion, character and context; preferred by ~75% of listeners (65-81% range) in blind head-to-head tests vs Cartesia Sonic 3.6, Inworld TTS-2, Gemini TTS, xAI TTS and GPT-4o mini TTS. (https://elevenlabs.io/blog/eleven-v4)
  - #1 on Artificial Analysis TTS arena: Took #1 on the Artificial Analysis Provider Voice TTS Arena (Elo ~1315-1319 at launch, ahead of Cartesia Sonic 3.6 at 1275 and Gemini 3.8 Flash TTS at 1267) and #1 on AA's Pronunciation Robustness benchmark, #2 on Controlled Voice. (https://artificialanalysis.ai/text-to-speech/leaderboard)
  - Real-time Turbo variant (~100 ms): eleven_v4_turbo: ~100 ms median inference latency, ~150 ms median time to first speech (vs Cartesia Sonic 3.6 262 ms, GPT-4o mini TTS 814 ms per ElevenLabs), for voice agents. (https://elevenlabs.io/v4)
  - Cross-lingual native accent, 90+ languages: 90+ languages (new: Cantonese, Mongolian, Odia); when target language differs from the reference voice, v4 speaks with a fluent native accent instead of carrying over the source accent. (https://elevenlabs.io/docs/overview/capabilities/text-to-speech/eleven-v4)
  - Inline tags incl. sound effects and free-text direction: Inline tags direct delivery, emotion, pacing, reactions, SFX and style, e.g. [laughs], [said angrily in French accent], [light rain], [phone buzzing], [quick, light, playful pace]. (https://elevenlabs.io/blog/eleven-v4)
  - Voice cloning from 10 s, PVC support restored: Instant Voice Clones from ~10 s of audio (docs still recommend 1-2 min); Professional Voice Clones supported again (not available on v3); speaker identity kept across regenerations/long-form. (https://elevenlabs.io/v4)
  - IPA pronunciation control: Pronunciation control with IPA support; more natural multi-speaker dialogue. (https://www.youtube.com/watch?v=th_tXR2QQ6U)

### Eleven Music v2.5
**Eleven Music v2.5** (ElevenLabs; current; music; released 2026-09-11) | ElevenLabs API: `music_v2_5` | Web app (ElevenMusic): https://elevenmusic.io — Announced 2026-09-11 (blog + YouTube). music_v2 and music_v1 remain available (v1 'outclassed by v2/v2.5'). Preferred over v2 in a blind test of 47,885 sample pairs; biggest gains in R&B/soul, hip hop/trap, rock/metal, orchestral/cinematic. Downloads: Free 5 lossless/day, Pro 400/month; tracks based on other artists' songs cannot be downloaded (protections built with labels/publishers).
  - ElevenLabs API: `music_v2_5` — https://api.elevenlabs.io/v1/music (docs: https://elevenlabs.io/docs/overview/capabilities/music)
  - Web app (ElevenMusic) — https://elevenmusic.io
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Commercially cleared music generation: Richer melodies and live-sounding instruments, built for commercial use; lossless downloads on every plan incl. Free. (https://elevenlabs.io/blog/music-v2-5-model)
  - Composition plans and audio reference: Music v2 line supports structured composition plans and reference-audio generation (v2.5 default for prompted and reference generation). (https://elevenlabs.io/docs/models)
  - Composition-plan chunks via API (found after launch): API support rolled out 2026-09-14 with 6,132-character composition chunks; waveform visual data via with_waveform_visual (2026-08-03). (https://elevenlabs.io/docs/changelog)

### Eleven v3 Conversational
**Eleven v3 Conversational** (ElevenLabs; current; audio/speech; released 2026-08-19) | ElevenLabs API: `eleven_v3_conversational` | ElevenAgents: https://elevenlabs.io/agents — GA announced 2026-08-19 (ElevenLabs X post and ElevenLabs Developers YouTube video). Artificial Analysis TTS arena Elo ~1196 (Aug 2026). Superseded for agents by eleven_v4_turbo (2026-09-28, ~100 ms). Exact streaming endpoint shown is the generic TTS stream endpoint; websockets also used in ElevenAgents.
  - ElevenLabs API: `eleven_v3_conversational` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream (docs: https://elevenlabs.io/docs/models)
  - ElevenAgents — https://elevenlabs.io/agents
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Real-time v3 with audio tags: Brings Eleven v3's expressive delivery and audio tags to streaming/real-time use at ~280 ms latency (excl. application & network), 70+ languages. (https://elevenlabs.io/docs/models)

### Scribe v2 / Scribe v2 Medical
**Scribe v2 / Scribe v2 Medical** (ElevenLabs; current; audio/speech; released 2026-01-09) | ElevenLabs API: `scribe_v2`; ElevenLabs API: `scribe_v2_medical` — Launched 2026-01-09; ElevenLabs claims 'the lowest word error rate recorded on industry-standard benchmarks' (FLEURS chart; company claim). Realtime variant in its own file. scribe_v1 (launched 2025-02-26, $0.40/h at launch) is deprecated ('outclassed by v2').
  - ElevenLabs API: `scribe_v2` — https://api.elevenlabs.io/v1/speech-to-text (docs: https://elevenlabs.io/docs/overview/capabilities/speech-to-text)
  - ElevenLabs API: `scribe_v2_medical` — https://api.elevenlabs.io/v1/speech-to-text (docs: https://elevenlabs.io/docs/models)
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Entity detection with timestamps: Native detection of PII, health and payment entities (56 categories at launch, 65 types per current docs) with exact timestamps. (https://elevenlabs.io/blog/introducing-scribe-v2)
  - Keyterm prompting, 32-speaker diarization: Keyterm prompting (100 terms at launch, now up to 1,000), speaker diarization up to 32 speakers, word timestamps, dynamic audio-event tagging, multi-language audio in one file; 90+ languages. (https://elevenlabs.io/docs/models)
  - Clinical variant (found after launch): scribe_v2_medical fine-tuned for clinical audio, HIPAA with BAA; generally available 2026-09-14. (https://elevenlabs.io/docs/changelog)

### Scribe v2 Realtime
**Scribe v2 Realtime** (ElevenLabs; current; audio/speech; released 2025-11-11) | ElevenLabs API (WebSocket): `scribe_v2_realtime` — Launched 2025-11-11; claims 93.5% accuracy across 30 European and Asian languages (company figure). EU and India data residency, zero-retention mode.
  - ElevenLabs API (WebSocket): `scribe_v2_realtime` — wss://api.elevenlabs.io/v1/speech-to-text/realtime (docs: https://elevenlabs.io/docs/models)
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - ~150 ms streaming STT with next-word prediction: Under 150 ms transcription latency with 'negative latency' next-word and punctuation prediction; VAD, manual commit, mid-conversation language switching; 90+ languages; PCM 48 kHz and u-law. (https://elevenlabs.io/blog/introducing-scribe-v2-realtime)
  - Realtime entity detection (found after launch): Entity detection added to realtime transcription on 2026-08-03. (https://elevenlabs.io/docs/changelog)

### Eleven Sound Effects v2
**Eleven Sound Effects v2** (ElevenLabs; current; audio/speech; released 2025-09) | ElevenLabs API: `eleven_text_to_sound_v2`; fal: `fal-ai/elevenlabs/sound-effects/v2` | Web app: https://elevenlabs.io/sound-effects — Release month (Sept 2025) is from third-party sources, not an official post. App pricing: 40 credits/second when duration is set.
  - ElevenLabs API: `eleven_text_to_sound_v2` — https://api.elevenlabs.io/v1/sound-generation (docs: https://elevenlabs.io/docs/overview/capabilities/sound-effects)
  - fal: `fal-ai/elevenlabs/sound-effects/v2` — https://fal.ai/models/fal-ai/elevenlabs/sound-effects/v2
  - Web app — https://elevenlabs.io/sound-effects
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Seamless looping SFX, 48 kHz: Text-to-sound effects up to 30 s per generation (0.1-30 s selectable), seamless looping for longer ambiences, prompt-influence control; MP3, WAV 48 kHz for non-looping. (https://elevenlabs.io/docs/overview/capabilities/sound-effects)

### Eleven v3
**Eleven v3** (ElevenLabs; current; audio/speech; released 2025-06-03) | ElevenLabs API: `eleven_v3`; ElevenLabs API (Text to Dialogue): `eleven_v3`; Runway API: `eleven_v3` | Web app: https://elevenlabs.io/app — Alpha announced 2025-06-03 (blog date); API initially via sales, GA across all platforms 2026-02-02. 70+ languages, 5,000 chars/request. Artificial Analysis TTS arena Elo ~1169 (Sept 2026). Professional Voice Clones not supported on v3 (restored in v4). Real-time variant eleven_v3_conversational has its own file. Voice design: eleven_ttv_v3. Superseded in quality by eleven_v4 (2026-09-28) but still current.
  - ElevenLabs API: `eleven_v3` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/models)
  - ElevenLabs API (Text to Dialogue): `eleven_v3` — https://api.elevenlabs.io/v1/text-to-dialogue
  - Runway API: `eleven_v3` — https://api.dev.runwayml.com/v1/text_to_speech (docs: https://docs.dev.runwayml.com/guides/models/)
  - Web app — https://elevenlabs.io/app
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Inline audio tags: Controls delivery with inline tags like [whispers], [laughs], [sighs], [excited]; marketed as 'the most expressive Text to Speech model' at launch. Not marked first: bracketed non-verbal cues existed earlier (e.g. Suno Bark, 2023). (https://elevenlabs.io/blog/eleven-v3)
  - Text to Dialogue (multi-speaker): Dedicated Text to Dialogue API for multi-speaker conversations with natural pacing and interruptions. (https://elevenlabs.io/blog/eleven-v3)
  - GA release with symbol/number normalization (found after launch): GA on 2026-02-02: preferred 72% of the time over alpha; error rate on numbers/symbols/notation cut 68% (15.3% -> 4.9%). (https://elevenlabs.io/blog/eleven-v3-is-now-generally-available)

### Eleven Flash v2.5 / Flash v2
**Eleven Flash v2.5 / Flash v2** (ElevenLabs; current; audio/speech; released 2024-12-18) | ElevenLabs API: `eleven_flash_v2_5`; ElevenLabs API (English only): `eleven_flash_v2` — Announced 2024-12-18 ('Meet Flash', X post). Replaced Turbo v2/v2.5 (now deprecated). Text normalization available for Flash v2.5 (enterprise). For expressive real-time use ElevenLabs now points to eleven_v4_turbo (~100 ms).
  - ElevenLabs API: `eleven_flash_v2_5` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/models)
  - ElevenLabs API (English only): `eleven_flash_v2` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/models)
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - ~75 ms TTS: Ultra-fast model for real-time use: ~75 ms model latency (excl. application & network). Flash v2.5: 32 languages, 40,000 chars/request; Flash v2: English only, 30,000 chars. (https://elevenlabs.io/docs/models)

### Eleven Multilingual v2
**Eleven Multilingual v2** (ElevenLabs; current; audio/speech; released 2023-08-22) | ElevenLabs API: `eleven_multilingual_v2` | Web app: https://elevenlabs.io/app — Launched 2023-08-22 (press date). Still current and the long-standing default for narration; supports style/speed settings and PVC. Superseded in expressiveness by v3/v4.
  - ElevenLabs API: `eleven_multilingual_v2` — https://api.elevenlabs.io/v1/text-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/models)
  - Web app — https://elevenlabs.io/app
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Stable long-form multilingual TTS: 'Lifelike model with rich emotional expression', 29 languages, 10,000 chars/request; keeps a voice's characteristics across languages. Launched as ElevenLabs exited beta. (https://elevenlabs.io/blog/elevenlabs-comes-out-of-beta-and-releases-eleven-multilingual-v2-a-foundational-ai-speech-model-for-nearly-30-languages)

### Eleven Multilingual STS v2 (Voice Changer)
**Eleven Multilingual STS v2 (Voice Changer)** (ElevenLabs; current; audio/speech) | ElevenLabs API: `eleven_multilingual_sts_v2`; ElevenLabs API (English only): `eleven_english_sts_v2` | Web app: https://elevenlabs.io/voice-changer — Release date not verified. Voice Isolator is also $0.12/min.
  - ElevenLabs API: `eleven_multilingual_sts_v2` — https://api.elevenlabs.io/v1/speech-to-speech/{voice_id} (docs: https://elevenlabs.io/docs/overview/capabilities/voice-changer)
  - ElevenLabs API (English only): `eleven_english_sts_v2` — https://api.elevenlabs.io/v1/speech-to-speech/{voice_id}
  - Web app — https://elevenlabs.io/voice-changer
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Speech-to-speech voice conversion: Converts a recording into another voice while keeping the original delivery (timing, emotion); multilingual model covers 29 languages. (https://elevenlabs.io/docs/models)

### Eleven Voice Design v3 (Text to Voice)
**Eleven Voice Design v3 (Text to Voice)** (ElevenLabs; current; audio/speech) | ElevenLabs API: `eleven_ttv_v3`; ElevenLabs API (older): `eleven_multilingual_ttv_v2` | Web app: https://elevenlabs.io/voice-design — Pricing not listed on the API pricing page (billed in credits). Release date not verified. ElevenLabs warns Voice Design voices may not perform as well on Eleven v4 as on earlier models.
  - ElevenLabs API: `eleven_ttv_v3` — https://api.elevenlabs.io/v1/text-to-voice/design (docs: https://elevenlabs.io/docs/models)
  - ElevenLabs API (older): `eleven_multilingual_ttv_v2` — https://api.elevenlabs.io/v1/text-to-voice/design
  - Web app — https://elevenlabs.io/voice-design
Notable capabilities:
  - Design a voice from a text description: Generates new synthetic voices from a prompt; eleven_ttv_v3 covers 70+ languages, eleven_multilingual_ttv_v2 29. (https://elevenlabs.io/docs/models)

### Eleven Dubbing v2
**Eleven Dubbing v2** (ElevenLabs; preview; audio/speech; released 2026-05-28) | Web app (ElevenCreative / ElevenProductions): https://elevenlabs.io/dubbing — Launched in UI 2026-05-28; API announced 2026-08-06 (blog) / changelog 2026-08-10. Docs label it 'Dubbing v2 Alpha' (default for Automatic Dubbing), hence status preview. No explicit model_id string found in API docs.
  - ElevenLabs API — https://api.elevenlabs.io/v1/dubbing (docs: https://elevenlabs.io/docs/overview/capabilities/dubbing)
  - Web app (ElevenCreative / ElevenProductions) — https://elevenlabs.io/dubbing
  - Pricing source: https://elevenlabs.io/pricing/api
Notable capabilities:
  - Direct speech-to-speech dubbing: Conditions directly on the original performance instead of an ASR -> translate -> TTS pipeline, so intonation and emotion carry across 90+ languages; ElevenLabs: 'For the first time, the emotion and performance of the original speaker carries across every language' (company claim, not independently verified as a first). (https://elevenlabs.io/blog/introducing-dubbing-v2)
  - Project-based dubbing API (found after launch): API (2026-08-06/10) with editable JSON transcripts/translations, regional variants (e.g. es-MX), sync-aware translation; 3 GB per file via API. (https://elevenlabs.io/blog/dubbing-api)

### Scribe v1
**Scribe v1** (ElevenLabs; deprecated; audio/speech; released 2025-02-26) | ElevenLabs API: `scribe_v1` — Deprecated on the models page ('First generation speech recognition (outclassed by v2)'). Use scribe_v2. Current price not listed separately.
  - ElevenLabs API: `scribe_v1` — https://api.elevenlabs.io/v1/speech-to-text (docs: https://elevenlabs.io/docs/models)
Notable capabilities:
  - ElevenLabs' first speech-to-text model: 99 languages, word timestamps, diarization and audio-event tagging; claimed highest benchmark accuracy vs Gemini 2.0 and Whisper v3 at launch ($0.40/hour). (https://elevenlabs.io/blog/meet-scribe)

### Eleven Turbo v2.5 / Turbo v2
**Eleven Turbo v2.5 / Turbo v2** (ElevenLabs; deprecated; audio/speech) | ElevenLabs API: `eleven_turbo_v2_5`; ElevenLabs API (English only): `eleven_turbo_v2` — Marked deprecated on the models page: 'First generation low-latency model (outclassed by Flash)'. Turbo v2.5: 32 languages; Turbo v2: English only. Migrate to eleven_flash_v2_5 or eleven_v4_turbo. Release dates (2024) not re-verified; shutdown date not stated.
  - ElevenLabs API: `eleven_turbo_v2_5` (docs: https://elevenlabs.io/docs/models)
  - ElevenLabs API (English only): `eleven_turbo_v2` (docs: https://elevenlabs.io/docs/models)
  - Pricing source: https://elevenlabs.io/pricing/api


## Figure AI

### Helix 2.5
**Helix 2.5** (Figure AI; current; robotics; released 2026-09-17) | None (runs only on Figure 03 robots; no public API, weights or waitlist): https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization — 'first' flags are Figure's 'to our knowledge' claims (first zero-shot whole-body generalization at this scope; first human-to-robot transfer scaling law measured on a humanoid). Company-reported results. Architecture/parameter counts not disclosed.
  - None (runs only on Figure 03 robots; no public API, weights or waitlist) — https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization
Notable capabilities:
  - [FIRST] Zero-shot whole-body generalization to unseen homes: 56% success (237/420 trials) tidying, towel folding and bed making in 30 never-seen Bay Area homes with no data from those homes; matched Helix 02's success with half the adaptation data. (https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization)
  - Pretrained from scratch on human video (Index): Pretrained from random initialization on Figure's Index human-video dataset (not a VLM); without it the same model scored 9%. (https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization)
  - [FIRST] Human-to-robot transfer scaling law: Predictable scaling of robot performance with human-video data (forecast error 0.54% over an 8x data range). (https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization)

### Helix 02
**Helix 02** (Figure AI; legacy; robotics; released 2026-01-27) | None (runs only on Figure 03 robots; no public API or weights): https://www.figure.ai/news/helix-02 — Figure: 'first demonstration of such long horizon, end-to-end pixels-to-whole body control on a humanoid robot' (company claim). Superseded by Helix 2.5 (2026-09-17). No external access.
  - None (runs only on Figure 03 robots; no public API or weights) — https://www.figure.ai/news/helix-02
Notable capabilities:
  - [FIRST] Pixels-to-whole-body control over long horizons: One visuomotor network links every sensor (vision, touch, proprioception) to every actuator; unloaded and reloaded a dishwasher across a full kitchen in a 4-minute run with walking, manipulation and balance, no resets. (https://www.figure.ai/news/helix-02)
  - System 0 learned whole-body controller: New 10M-parameter S0 at 1 kHz trained on 1,000+ hours of retargeted human motion and 200,000+ parallel simulated environments, under S1 (200 Hz) and S2 (semantic reasoning). (https://www.figure.ai/news/helix-02)
  - Tactile and palm-camera policies: First Figure policies that depend on Figure 03's palm cameras and fingertip tactile sensing for occluded, delicate manipulation. (https://www.figure.ai/news/helix-02)

### Helix (Figure, v1)
**Helix (Figure, v1)** (Figure AI; legacy; robotics; released 2025-02-20) | None (runs only on Figure robots; no public API or weights): https://www.figure.ai/news/helix — 'first' flags are Figure's own claims at announcement (2025-02-20). Superseded by Helix 02 (2026-01) and Helix 2.5 (2026-09). Never publicly available.
  - None (runs only on Figure robots; no public API or weights) — https://www.figure.ai/news/helix
Notable capabilities:
  - [FIRST] Full upper-body humanoid control from a VLA: Continuous high-rate control of the whole humanoid upper body (wrists, torso, head, individual fingers) over a 35-DoF action space. (https://www.figure.ai/news/helix)
  - Dual-system architecture (S2 + S1): System 2: 7B VLM at 7-9 Hz for scene/language understanding; System 1: 80M-parameter visuomotor transformer at 200 Hz. (https://www.figure.ai/news/helix)
  - [FIRST] Multi-robot collaboration with one set of weights: Same model ran simultaneously on two robots collaborating on a shared grocery-storage task. (https://www.figure.ai/news/helix)
  - [FIRST] Fully onboard on embedded low-power GPUs: Runs entirely on the robot's embedded GPUs - Figure calls it the first VLA ready for commercial deployment this way. (https://www.figure.ai/news/helix)


## Fish Audio

### Fish Audio S2 Pro / S2.1 Pro
**Fish Audio S2 Pro / S2.1 Pro** (Fish Audio; current; audio/speech; released 2026-03-09; open weights) | Fish Audio API: `s2.1-pro`; Fish Audio API (free tier): `s2.1-pro-free`; Fish Audio API: `s2-pro`; Hugging Face: `fishaudio/s2-pro` | OpenRouter: https://openrouter.ai/fish-audio/s2.1-pro; GitHub: https://github.com/fishaudio/fish-speech — S2 Pro held #1 open-weights on Artificial Analysis until Breeze TTS 2 (Aug 2026); now #2 open (~1119 Elo). S2.1 Pro weights are NOT open. OpenRouter lists S2.1 Pro release as 2026-07-29 (API availability there). Predecessor OpenAudio S1 (`s1`) still supported.
  - Fish Audio API: `s2.1-pro` — https://api.fish.audio/v1/tts (model passed in `model` header) (docs: https://docs.fish.audio/developer-guide/getting-started/changelog)
  - Fish Audio API (free tier): `s2.1-pro-free` — https://api.fish.audio/v1/tts
  - Fish Audio API: `s2-pro`
  - Hugging Face: `fishaudio/s2-pro` — https://huggingface.co/fishaudio/s2-pro
  - OpenRouter — https://openrouter.ai/fish-audio/s2.1-pro
  - GitHub — https://github.com/fishaudio/fish-speech
  - Pricing source: https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits
Notable capabilities:
  - Inline natural-language emotion/paralinguistic tags: Free-form bracket cues like [whisper], [laugh], [emphasis]; multi-speaker dialogue in one pass; 80+ languages from 10M+ hours of training audio. (https://fish.audio/blog/fish-audio-open-sources-s2/)
  - Open model with production inference stack: Dual-AR (4B slow + 400M fast) on a Qwen3-4B backbone released with fine-tuning code and SGLang serving; RTF 0.195, ~100 ms TTFA; Seed-TTS Eval WER 0.54% zh / 0.99% en. (https://arxiv.org/abs/2603.08823)
  - Free production API (S2.1 Pro) (found after launch): S2.1 Pro (closed, 2026-06-23) offered free under fair use with ~90 ms TTFA, 83 languages; 61% win rate vs S2 Pro. (https://fish.audio/blog/s2-1-pro-free-api/)


## Generalist AI

### Generalist GEN-1.5
**Generalist GEN-1.5** (Generalist AI; current; robotics; released 2026-08-19) | Generalist AI partners (no public access announced): https://generalistai.com/blog/gen-1.5 — Released 6 days before Skild S1, which makes a similar one-video in-context claim for long-horizon tasks. Company-reported. Video: https://www.youtube.com/watch?v=1cllCVK-9lo
  - Generalist AI partners (no public access announced) — https://generalistai.com/blog/gen-1.5
Notable capabilities:
  - [FIRST] One-shot learning of dexterous closed-loop tasks: Learns new tasks in-context from one demonstration video: 59% average success one-shot across 10 tasks; 83% with few-shot adaptation (10 gradient steps on 5 minutes of data). Generalist says it is the first model it knows of to show this across a wide range of dexterous closed-loop tasks. (https://generalistai.com/blog/gen-1.5)
  - 30-second video memory, 100 Hz actions: Takes video (30 s memory window), sensors, language and proprioception and outputs 100 Hz action trajectories. (https://generalistai.com/blog/gen-1.5)

### Generalist GEN-1
**Generalist GEN-1** (Generalist AI; current; robotics; released 2026-04-02) | Generalist AI early-access partners (partnerships@generalistai.com): https://generalistai.com/blog/gen-1 — No public API/weights; early-access partners only. Successor GEN-1.5 (2026-08-19) adds one-shot learning (see generalist-gen-1-5). Results are company-reported. Video: https://www.youtube.com/watch?v=SY2xyrmV44Y
  - Generalist AI early-access partners (partnerships@generalistai.com) — https://generalistai.com/blog/gen-1
Notable capabilities:
  - [FIRST] Mastery of simple physical tasks: 99% success on several tasks (GEN-0: 64%), ~3x faster than prior state of the art, ~1 hour of robot data per task; Generalist calls it the first general-purpose model to cross a 'mastery' threshold for simple tasks. (https://generalistai.com/blog/gen-1)
  - Pretrained on 500k+ hours of human wearable data: Pretraining dataset of 500,000+ hours of real-world physical interaction captured with wearable devices on humans (no robot data), spanning many end effectors; later extended to a broad range of end effectors from five-finger hands to special tools. (https://generalistai.com/blog/gen-1)
  - Robotics scaling laws (GEN-0 predecessor): GEN-0 (Nov 2025) showed scaling laws for robot foundation models, with all tracked zero-shot tasks improving together as pretraining scaled. (https://generalistai.com/blog/gen-1)


## Google

### Chirp 3 Transcription (Google Cloud Speech-to-Text)
**Chirp 3 Transcription (Google Cloud Speech-to-Text)** (Google; current; audio/speech; released 2025-10-13) | Google Cloud Speech-to-Text API V2: `chirp_3` — Private preview 2025-04-11, public preview 2025-08-29, GA 2025-10-13 (US/EU multi-region). No word-level timestamps or word confidence. For developers, Gemini 3.5 Transcribe (2026-08-26) claims 70% faster time-to-final than Chirp 3.
  - Google Cloud Speech-to-Text API V2: `chirp_3` — https://speech.googleapis.com/v2 (Recognize, StreamingRecognize, BatchRecognize) (docs: https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3)
  - Pricing source: https://cloud.google.com/speech-to-text/pricing
Notable capabilities:
  - Multilingual ASR with language-agnostic mode: ~100+ languages/locales (about 20 GA), language_codes=['auto'] for language-agnostic transcription, diarization in ~15 languages, speech adaptation. (https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3)

### Chirp 3 HD voices (Google Cloud Text-to-Speech)
**Chirp 3 HD voices (Google Cloud Text-to-Speech)** (Google; current; audio/speech; released 2025-04-02) | Google Cloud Text-to-Speech API: `<locale>-Chirp3-HD-<Voice> (e.g. en-US-Chirp3-HD-Charon)` — GA 2025-04-02 (8 speakers, 31 locales), since expanded to 60+ locales. Google's enterprise, non-LLM TTS line; the Gemini-TTS models (gemini-2.5-*-tts, Gemini 3.1 Flash TTS) are offered in the same Cloud TTS API. No 2026 successor (e.g. 'Chirp 4') found.
  - Google Cloud Text-to-Speech API: `<locale>-Chirp3-HD-<Voice> (e.g. en-US-Chirp3-HD-Charon)` — https://texttospeech.googleapis.com (regions global, us, eu, asia-southeast1, europe-west2, asia-northeast1) (docs: https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd)
  - Pricing source: https://cloud.google.com/text-to-speech/pricing
Notable capabilities:
  - Streaming HD voices in 60+ locales: 28 named voices, streaming and batch synthesis, pace (0.25-2x), pause and IPA/X-SAMPA pronunciation controls, SSML. (https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd)
  - Instant custom voice (found after launch): Chirp 3 Instant custom voice clones a voice from a short sample (30+ locales), priced at $60 per 1M characters. (https://docs.cloud.google.com/text-to-speech/docs/release-notes)


## Google DeepMind

### Gemini 3.8 Flash TTS
**Gemini 3.8 Flash TTS** (Google DeepMind; current; audio/speech; released 2026-09-22) | ctx 8,192 | $0.5 in / $9 out per 1M tokens (text in / audio out; introductory through 2026-12-31, $1.00 / $18.00 from 2027-01-01) | Gemini API: `gemini-3.8-flash-tts`; Gemini API: `gemini-3.8-flash-lite-tts` — Sibling gemini-3.8-flash-lite-tts (101 languages) costs $0.50 in / $6.00 audio out (intro). Outputs SynthID-watermarked. Older: gemini-3.1-flash-tts-preview, gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts.
  - Gemini API: `gemini-3.8-flash-tts` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash-tts:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts)
  - Gemini API: `gemini-3.8-flash-lite-tts` (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-lite-tts)
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Voice design from prompts: Create entirely new voices from natural-language descriptions; #1 on Hume AI Voice Design Benchmark. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/)
  - #1 on Hume Real-World VoiceEQ leaderboard (found after launch): Hume's blind human-rated benchmark (2026-09-24): Gemini 3.8 Flash TTS 0.920 and Flash-Lite TTS 0.914 expressivity-reliability score, ahead of Gemini 2.5 Pro TTS (0.880) and Cartesia Sonic 3.6 (0.840); long-form stability up from 1.22 to ~2.9-3.0/5, but weaker speaker similarity (3.68/5). (https://www.hume.ai/blog/newly-released-google-s-gemini-3-8-flash-tts-tops-hume-s-real-world-voiceeq-leaderboard)
  - Voice replication: Recreates a consistent voice from a ~30-second sample with consent verification; 2,000+ library voices. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/)
  - Directed long-form multi-speaker audio: Line-by-line direction of pacing/emotion, dual-speaker staging, stable over hours; 130+ languages with auto-detection. (https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts)

### Gemini 3.8 Live
**Gemini 3.8 Live** (Google DeepMind; current; audio/speech; released 2026-09-15) | ctx 131,072 | $0.75 in / $4.5 out per 1M tokens (text in $0.75, text out $4.50; audio in $3.00 = ~$0.005/min, audio out $12.00 = ~$0.018/min) | Gemini Live API (WebSocket): `gemini-3.8-live`; Gemini Live API (WebSocket): `gemini-3.8-live-extended-thinking` | Web app: https://gemini.google.com — Default Live API model; thinking_level not supported on gemini-3.8-live (use gemini-3.8-live-extended-thinking for deeper reasoning; pricing page lists it at the same rates as gemini-3.8-live, checked 2026-09-29). Previous: gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025. WebSocket endpoint is the standard Live API URL, not re-read today.
  - Gemini Live API (WebSocket): `gemini-3.8-live` — wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live)
  - Gemini Live API (WebSocket): `gemini-3.8-live-extended-thinking` (docs: https://ai.google.dev/gemini-api/docs/live-api)
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Real-time multilingual voice agents: Low-latency speech-to-speech with near-real-time visual grounding; 97 languages with mid-conversation switching. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/)
  - Asynchronous tool use while talking: Keeps the conversation going while tools run in the background, narrating progress ('Let me check that...'). (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/)
  - Extended Thinking variant tops S2S quality: gemini-3.8-live-extended-thinking ranked #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6) and 97.7% Big Bench Audio. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/)

### Gemini 3.8 Flash
**Gemini 3.8 Flash** (Google DeepMind; current; reasoning-llm; released 2026-09-02) | ctx 1,048,576 | $0.75 in / $3.75 out per 1M tokens (Standard tier, introductory price through 2026-12-31; rises to $1.50 / $7.50 from 2027-01-01) | Gemini API: `gemini-3.8-flash`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-3.8-flash`; OpenRouter: `google/gemini-3.8-flash` | Web app: https://gemini.google.com — Newest and recommended Gemini text model as of 2026-09 (no Pro newer than 3.1 Pro preview; 3.5 Pro announced but unreleased). Aliases gemini-flash-latest may point here. Model card says some domains' knowledge only to 2025-01. Vertex id inferred from docs page.
  - Gemini API: `gemini-3.8-flash` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-3.8-flash` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash)
  - OpenRouter: `google/gemini-3.8-flash`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Long-horizon software engineering: Google's most capable Flash for autonomous end-to-end engineering; 73.7% on DeepSWE v1.1, 89.4% terminal-based coding per model card. (https://deepmind.google/models/model-cards/gemini-3-8-flash/)
  - Specialized-domain agentic analysis: Beats 3.7 Flash and other frontier models on Vals Finance Agent v2 (61.4%) and Harvey's Legal Agent Benchmark. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/)
  - Agentic long-video understanding: 87.8% long video understanding in agentic mode (agentic video understanding added for 3.x Flash on 2026-09-01). (https://deepmind.google/models/model-cards/gemini-3-8-flash/)
  - Adjustable thinking levels + computer use: Thinking low/medium/high, computer use (preview), Maps/Search grounding, flex and priority inference tiers. (https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash)
  - Cyber sibling model: Launched alongside Gemini 3.8 Flash Cyber (vulnerability detection/patching, 47.2% CWE-Bench pass@1), available only to vetted defenders via the Fairwind Program. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/)

### Gemini 3.5 Transcribe (and Transcribe Live)
**Gemini 3.5 Transcribe (and Transcribe Live)** (Google DeepMind; current; audio/speech; released 2026-08-26) | $? in / $12 out per 1M tokens (USD) for gemini-3.5-transcribe (~$0.003/min audio in + ~$0.002/min text out); gemini-3.5-transcribe-live $3.50 in / $21.00 out (~$0.005 + ~$0.004 per min) | Gemini API (Interactions API, files): `gemini-3.5-transcribe`; Gemini Live API (WebSocket streaming): `gemini-3.5-transcribe-live` | Google AI Studio: https://aistudio.google.com — Changelog lists both ids GA on 2026-08-26, while the launch blog says public preview in AI Studio and Gemini Enterprise Agent Platform. Limits: 1 h per file request (30 min with diarization/timestamps), 10 min per live session. Diarization: docs say up to 8 speakers, blog says up to three - unresolved. Powers Rambler on Android and the Gemini app on macOS; coming to Chrome and Gboard. Press quotes $0.005/min (file) and $0.009/min (live) all-in. Model card (read 2026-09-29): https://deepmind.google/models/model-cards/gemini-3-5-audio/ - covers Gemini 3.5 Live Translate, Transcribe and Transcribe Live (card dated 2026-08-26); no numeric evals in the card itself; knowledge cutoff January 2025; did not reach any Tracked or Critical Capability Levels under the Frontier Safety Framework. Card lists hallucinations and occasional slowness/timeouts as limitations; surfaces: Antigravity, Gboard, Gemini app, Vertex AI, Google Workspace.
  - Gemini API (Interactions API, files): `gemini-3.5-transcribe` (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe)
  - Gemini Live API (WebSocket streaming): `gemini-3.5-transcribe-live` (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe)
  - Google AI Studio — https://aistudio.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Smart transcription: Handles self-corrections, removes filler words and auto-formats text; custom vocabulary biasing up to 1,000 terms. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/)
  - Low word error rate: Google cites Artificial Analysis WER of 2.6% (non-streaming) and 4.0% (streaming); 70% faster time-to-final than Chirp 3. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/)
  - 85+ languages with code-switching, diarization, word timestamps: Utterance-level language detection across 85+ languages; speaker diarization; word-level timestamps (not combinable with custom vocabulary). (https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe)

### Lyria 3.5
**Lyria 3.5** (Google DeepMind; current; music; released 2026-07-29) | Gemini API (Interactions API): `lyria-3.5`; Gemini API (Interactions API): `lyria-3-clip-preview`; OpenRouter: `google/lyria-3-pro-preview` | Web app: https://gemini.google.com — Launched 2026-07-29 in Google Flow Music (the rebranded ProducerAI); Gemini API GA 2026-09-03 (status Stable, no free tier). Not yet listed on the Vertex/Agent Platform Lyria pages or pricing as of 2026-09-29 (Vertex still offers lyria-3-pro-preview, lyria-3-clip-preview and lyria-002). The lyria-3-clip-preview access line above is the older Lyria 3 Clip, see lyria-3.md; lyria-realtime-exp covers streaming music (lyria-realtime.md). OpenRouter lists only Lyria 3 previews (not 3.5).
  - Gemini API (Interactions API): `lyria-3.5` — https://generativelanguage.googleapis.com/v1beta/interactions (docs: https://ai.google.dev/gemini-api/docs/music-generation)
  - Gemini API (Interactions API): `lyria-3-clip-preview` — https://generativelanguage.googleapis.com/v1beta/interactions
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3)
  - OpenRouter: `google/lyria-3-pro-preview`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Full songs with vocals and lyrics: Full-length ~2-minute tracks with verses/choruses/bridges, generated vocals and lyrics; 44.1 kHz stereo MP3/WAV. (https://ai.google.dev/gemini-api/docs/music-generation)
  - Image-conditioned music: Accepts text and image prompts via the Interactions API. (https://ai.google.dev/gemini-api/docs/music-generation)
  - SynthID-watermarked audio: Latent-diffusion model with SynthID watermarking on outputs. (https://deepmind.google/models/model-cards/lyria-3-5/)

### Gemini 3.5 Flash-Lite
**Gemini 3.5 Flash-Lite** (Google DeepMind; current; llm; released 2026-07-21) | ctx 1,048,576 | $0.3 in / $2.5 out per 1M tokens (Standard tier) | Gemini API: `gemini-3.5-flash-lite`; OpenRouter: `google/gemini-3.5-flash-lite` | Web app: https://gemini.google.com — Cheapest current Gemini text model; recommended replacement for 2.5 Flash/Flash-Lite and 3.1 Flash-Lite. Alias gemini-flash-lite-latest may point here (not verified).
  - Gemini API: `gemini-3.5-flash-lite` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite)
  - OpenRouter: `google/gemini-3.5-flash-lite`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - High-throughput subagent model: ~350 output tokens/s; optimized for subagent tasks and document processing at low cost. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)
  - Strong coding for a Lite tier: 54% on Terminal-Bench 2.1 vs 31% for the previous Flash-Lite. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)

### Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image)
**Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image)** (Google DeepMind; current; image-gen; released 2026-06-30) | $0.25 in / $1.5 out per 1M tokens text/image/video input and text output; image output $30 per 1M tokens (~$0.0336 per 1K image) | Gemini API: `gemini-3.1-flash-lite-image`; OpenRouter: `google/gemini-3.1-flash-lite-image` — GA in Gemini API 2026-06-30 per changelog. Token limits not verified on docs (OpenRouter lists 65,536 context).
  - Gemini API: `gemini-3.1-flash-lite-image` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite-image:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite-image)
  - OpenRouter: `google/gemini-3.1-flash-lite-image`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Lowest-cost Gemini image model: About half the per-image price of Nano Banana 2 (~$0.034 per 1K image). (https://ai.google.dev/gemini-api/docs/pricing)
  - Video-as-input image generation: Accepts text, image and video inputs for image generation/editing. (https://ai.google.dev/gemini-api/docs/pricing)

### Gemini Omni Flash (Omni 1.1 Flash)
**Gemini Omni Flash (Omni 1.1 Flash)** (Google DeepMind; current; video-gen; released 2026-05) | ctx 1,048,576 | $1.5 in / $9 out per 1M tokens for text/image/video/audio input ($1.50) and text output ($9.00); video output $17.50 per 1M tokens (~$0.10 per second at 720p) | Gemini API (Interactions API): `gemini-omni-1.1-flash` | Web app: https://gemini.google.com — Announced at Google I/O 2026 (preview id gemini-omni-flash-preview); Omni 1.1 Flash GA in the API 2026-08-27. Google's recommended default video model over Veo 3.1. Live API model list reports 131k context for gemini-omni-1.1-flash vs 1M on the docs page. Exact I/O day not verified.
  - Gemini API (Interactions API): `gemini-omni-1.1-flash` — https://generativelanguage.googleapis.com/v1beta/interactions (docs: https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/omni-1-1-flash)
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Any-input video generation: Generates video with native audio from any mix of text, image, audio and video input, grounded in Gemini world knowledge. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/)
  - Conversational video editing: Edit, extend (inputs up to 10 s), interpolate keyframes and upscale videos through multi-turn natural-language conversation. (https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash)
  - Up to 4K output: 3-10 s clips at 360p/720p/1080p/4K, 24 fps. (https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash)
  - Avatars + SynthID: Launched with avatar support (your own digital likeness); all outputs carry SynthID watermarks. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/)

### Gemini Embedding 2
**Gemini Embedding 2** (Google DeepMind; current; embedding; released 2026-04-22) | ctx 8,192 | $0.2 in / $? out per 1M text input tokens; image $0.45/1M (~$0.00012 per image), audio $6.50/1M (~$0.00016/s), video $12.00/1M (~$0.00079 per frame) | Gemini API: `gemini-embedding-2` — Public preview March 2026 (id gemini-embedding-2-preview, still listed), GA 2026-04-22. modality_out 'text' is a placeholder: output is a vector. Predecessor gemini-embedding-001 (text-only) shuts down 2028-05-14.
  - Gemini API: `gemini-embedding-2` — https://generativelanguage.googleapis.com/v1beta/models/gemini-embedding-2:embedContent (docs: https://ai.google.dev/gemini-api/docs/embeddings)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/embedding-2)
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Natively multimodal embeddings: Text, images, video, audio and PDFs mapped into one embedding space; Google's first natively multimodal embedding model and first in the Gemini API. (https://ai.google.dev/gemini-api/docs/embeddings)
  - Matryoshka dimensions: Flexible 128-3072 output dimensions (recommended 768/1536/3072); 100+ languages. (https://ai.google.dev/gemini-api/docs/embeddings)

### Gemma 4
**Gemma 4** (Google DeepMind; current; llm; released 2026-04-02; open weights) | ctx 262,144 | $0.09 in / $0.34 out per 1M tokens, OpenRouter price for google/gemma-4-31b-it (26B-A4B: $0.09 / $0.30; free variants exist). Weights free to download | Hugging Face: `google/gemma-4-31B-it`; Hugging Face: `google/gemma-4-26B-A4B-it`; Hugging Face: `google/gemma-4-12B-it`; Hugging Face: `google/gemma-4-E4B-it`; Hugging Face: `google/gemma-4-E2B-it`; OpenRouter: `google/gemma-4-31b-it`; OpenRouter: `google/gemma-4-26b-a4b-it` — Sizes E2B, E4B (128K context), 12B, 26B A4B MoE, 31B dense (256K context); base and -it variants plus QAT/GGUF quantized repos. 12B released later (HF repo 2026-05-23). Audio input on E2B/E4B/12B only. 140+ languages. pricing is third-party (OpenRouter), not Google.
  - Hugging Face: `google/gemma-4-31B-it` — https://huggingface.co/google/gemma-4-31B-it
  - Hugging Face: `google/gemma-4-26B-A4B-it` — https://huggingface.co/google/gemma-4-26B-A4B-it
  - Hugging Face: `google/gemma-4-12B-it` — https://huggingface.co/google/gemma-4-12B-it
  - Hugging Face: `google/gemma-4-E4B-it` — https://huggingface.co/google/gemma-4-E4B-it
  - Hugging Face: `google/gemma-4-E2B-it` — https://huggingface.co/google/gemma-4-E2B-it
  - OpenRouter: `google/gemma-4-31b-it`
  - OpenRouter: `google/gemma-4-26b-a4b-it`
  - Google docs (docs: https://ai.google.dev/gemma/docs/core)
  - Pricing source: https://openrouter.ai/google/gemma-4-31b-it
Notable capabilities:
  - First Apache-2.0 Gemma: First Gemma generation under the permissive Apache 2.0 license instead of Google's custom Gemma terms. (https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)
  - Intelligence per parameter: 31B dense ranked #3 and 26B A4B MoE #6 among open models on Arena at launch. (https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)
  - On-device agentic models: E2B/E4B edge models with native audio+vision, function calling and structured JSON, running offline on phones/Raspberry Pi/Jetson. (https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)
  - Encoder-free unified 12B (found after launch): Gemma 4 12B, added later, is a unified encoder-free multimodal model with native audio. (https://ai.google.dev/gemma/docs/core)

### Nano Banana 2 (Gemini 3.1 Flash Image)
**Nano Banana 2 (Gemini 3.1 Flash Image)** (Google DeepMind; current; image-gen; released 2026-02-26) | ctx 131,072 | $0.5 in / $3 out per 1M tokens text/image input and text output; image output $60 per 1M tokens = $0.045 (0.5K) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) per image | Gemini API: `gemini-3.1-flash-image`; OpenRouter: `google/gemini-3.1-flash-image` | Web app: https://gemini.google.com — Preview id gemini-3.1-flash-image-preview (2026-02-26, still served); stable id GA 2026-05-28. Replacement for gemini-2.5-flash-image and Imagen 4. OpenRouter lists 131k context for stable id.
  - Gemini API: `gemini-3.1-flash-image` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-flash-image)
  - OpenRouter: `google/gemini-3.1-flash-image`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Pro quality at Flash speed: Brings Nano Banana Pro world knowledge, reasoning and quality to a fast Flash model. (https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/)
  - Image search grounding: Uses real-time web/image search to render real subjects accurately; supports thinking. (https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image)
  - Text rendering and in-image translation: Legible text for marketing assets and translation of text inside images. (https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/)
  - Extreme aspect ratios and 512px-4K: 0.5K/1K/2K/4K outputs and 1:4, 4:1, 1:8, 8:1 ratios; consistency of up to 5 characters and 14 objects. (https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image)

### Nano Banana Pro (Gemini 3 Pro Image)
**Nano Banana Pro (Gemini 3 Pro Image)** (Google DeepMind; current; image-gen; released 2025-11-20) | ctx 65,536 | $2 in / $12 out per 1M tokens text/image input and text output; image output $120 per 1M tokens = $0.134 per 1K/2K image, $0.24 per 4K image | Gemini API: `gemini-3-pro-image`; OpenRouter: `google/gemini-3-pro-image` | Web app: https://gemini.google.com — Launched 2025-11-20 as gemini-3-pro-image-preview (still on OpenRouter); stable id GA 2026-05-28. Highest-quality but priciest Gemini image model.
  - Gemini API: `gemini-3-pro-image` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-image)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-pro-image)
  - OpenRouter: `google/gemini-3-pro-image`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Accurate multilingual text in images: Correct, legible text rendering in many languages, fonts and calligraphy; suited to infographics and mockups. (https://blog.google/technology/ai/nano-banana-pro/)
  - Search-grounded visuals: Uses Google Search to visualize real-time info (weather, sports, recipes) and factual data visualizations. (https://blog.google/technology/ai/nano-banana-pro/)
  - Multi-image composition: Blends up to 14 images while keeping resemblance of up to 5 people; up to 4K with lighting/depth-of-field edits. (https://blog.google/technology/ai/nano-banana-pro/)

### Gemini Robotics 2
**Gemini Robotics 2** (Google DeepMind; preview; robotics; released 2026-07-30) | Gemini Robotics trusted tester / early-access program (application form): https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform — Vision-language-action model (outputs robot motor commands; modality 'action'). No public API or weights: available only to early-access partners (Apptronik, Boston Dynamics, Agile Robots, Franka, 100+ trusted testers) via waitlist form. DeepMind says 'for the first time, our model can control entire humanoid robots' - first for Google, not industry-first (Figure Helix 02 showed whole-body VLA control in Jan 2026). Predecessor: Gemini Robotics 1.5 (Sep 2025), itself trusted-tester only.
  - Gemini Robotics trusted tester / early-access program (application form) — https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform (docs: https://deepmind.google/models/gemini-robotics/)
Notable capabilities:
  - Whole-body humanoid control from a VLA: Google's first VLA to control an entire humanoid (walking, crouching, balancing while manipulating) rather than only the upper body; e.g. Apollo with Inspire hands: 68.4% pick from table, 45.7% from floor, 76.3% from shelf. (https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)
  - Multi-finger and gripper dexterity across embodiments: Same model drives multi-fingered hands and grippers (Franka Duo: 89.6% precise insertion; Apollo with SharpaWave hands: 92% unscrew bulb). (https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)
  - Paired with ER 2 planner: Designed to be called by Gemini Robotics ER 2, which plans, tracks progress and coordinates multiple robots. (https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)

### Gemini Robotics ER 2
**Gemini Robotics ER 2** (Google DeepMind; preview; robotics; released 2026-07-30) | ctx 131,072 | $1 in / $5 out per 1M tokens (text/image/video/audio input); introductory rate through 2026-12-31, rising to $2.00 in / $10.00 out from 2027-01-01; Batch API half price | Gemini API: `gemini-robotics-er-2-preview`; Gemini API (Live API, streaming): `gemini-robotics-er-2-streaming-preview` | Google AI Studio: https://aistudio.google.com; Sample code (GitHub): https://github.com/google-gemini/robotics-samples — Vision-language model for robotics (outputs text/JSON, not motor commands). 131,072 input / 65,536 output tokens. Standard id supports caching, code execution, computer use, file search, function calling, Search and Maps grounding, structured outputs and thinking. Replaces gemini-robotics-er-1.6-preview (shut down 2026-08-31). No GA id yet. Knowledge cutoff not stated.
  - Gemini API: `gemini-robotics-er-2-preview` — https://generativelanguage.googleapis.com/v1beta/models/gemini-robotics-er-2-preview:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview)
  - Gemini API (Live API, streaming): `gemini-robotics-er-2-streaming-preview` (docs: https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-streaming-preview)
  - Google AI Studio — https://aistudio.google.com
  - Gemini Enterprise Agent Platform (Google Cloud, private preview) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/gemini-robotics-er)
  - Sample code (GitHub) — https://github.com/google-gemini/robotics-samples
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Embodied reasoning "robot brain" in a public API: Spatial reasoning (points, boxes, trajectories), multi-step task planning, tool/function calling and code execution to orchestrate a robot's VLA or controller; publicly callable, unlike the VLA models. (https://ai.google.dev/gemini-api/docs/robotics-overview)
  - Continuous video monitoring and task-progress tracking: Watches video feeds to track progress and adapt; Google reports 91.3% moment-finding accuracy (0.96 s mean absolute distance) at ~4x the speed of the previous generation and 57.4% progress classification. (https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/)
  - Low-latency streaming via Live API: Separate gemini-robotics-er-2-streaming-preview id supports bidirectional audio/video streaming with function calling and thinking (no caching, code execution or structured output). (https://ai.google.dev/gemini-api/docs/robotics-streaming)
  - Multi-robot collaboration: Coordinates heterogeneous robots (e.g. wheeled rovers and humanoids, Boston Dynamics Spot demo) to communicate and hand off tasks. (https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/)

### Gemini Robotics On-Device 2
**Gemini Robotics On-Device 2** (Google DeepMind; preview; robotics; released 2026-07-30) | Gemini Robotics trusted tester / early-access program (application form): https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform — Successor to Gemini Robotics On-Device (June 2025). Trusted-tester / partner access only; parameter count and hardware requirements not published.
  - Gemini Robotics trusted tester / early-access program (application form) — https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform (docs: https://deepmind.google/models/gemini-robotics/)
Notable capabilities:
  - Local VLA inference on robot hardware: Lightweight version of the Gemini Robotics VLA optimized to run locally without a network connection. (https://deepmind.google/models/gemini-robotics/)
  - Fast adaptation to new embodiments: Adapts to completely new robot bodies with a few hours of data; typically fewer than 200 examples for a new bi-arm robot (uses motion transfer from Gemini Robotics 1.5). (https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)

### Gemini 3.5 Live Translate
**Gemini 3.5 Live Translate** (Google DeepMind; preview; audio/speech; released 2026-06-09) | ctx 131,072 | Gemini Live API (WebSocket): `gemini-3.5-live-translate-preview` | Google AI Studio: https://aistudio.google.com/live?model=gemini-3.5-live-translate-preview; Google Translate app / Google Meet: https://translate.google.com — Public preview in the Live API/AI Studio from 2026-06-09; Meet private preview; Google Translate on Android/iOS (incl. headphone 'listening mode'). Outputs SynthID-watermarked. No function calling, thinking or caching. Model card (read 2026-09-29): https://deepmind.google/models/model-cards/gemini-3-5-audio/ - covers Gemini 3.5 Live Translate, Transcribe and Transcribe Live (card dated 2026-08-26); no numeric evals in the card itself; knowledge cutoff January 2025; did not reach any Tracked or Critical Capability Levels under the Frontier Safety Framework. Card-listed Live Translate limitations: inconsistent voices, language detection struggles with non-native accents and rapid switching, imperfect background-noise handling, occasional audio artifacts. OpenAI's rival gpt-realtime-translate launched a month earlier (2026-05-07).
  - Gemini Live API (WebSocket): `gemini-3.5-live-translate-preview` (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-live-translate-preview)
  - Google AI Studio — https://aistudio.google.com/live?model=gemini-3.5-live-translate-preview
  - Google Translate app / Google Meet — https://translate.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Continuous speech-to-speech translation preserving the speaker's voice: Audio-to-audio (no ASR-MT-TTS cascade), generating speech continuously a few seconds behind the speaker while keeping intonation, pacing and pitch. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/)
  - 70+ languages, 2,000+ pairs: Auto-detects 70+ languages and supports 2,000+ language combinations in one meeting; expands Google Meet live translation from 5 to 70+ languages. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/)

### Gemini 3.1 Pro
**Gemini 3.1 Pro** (Google DeepMind; preview; reasoning-llm; released 2026-02-19) | ctx 1,048,576 | $2 in / $12 out per 1M tokens (Standard, prompts <=200k; >200k: $4.00 in / $18.00 out) | Gemini API: `gemini-3.1-pro-preview`; OpenRouter: `google/gemini-3.1-pro-preview` | Web app: https://gemini.google.com — Still the newest Pro model in the Gemini API (preview only; Gemini 3.5 Pro announced at I/O 2026 but not released as of 2026-09). Predecessor gemini-3-pro-preview is shut down. Newer 3.5+ Flash models beat it on many agentic/coding benchmarks at lower cost. Vertex id not verified.
  - Gemini API: `gemini-3.1-pro-preview` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-pro)
  - OpenRouter: `google/gemini-3.1-pro-preview`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Novel-pattern reasoning: Verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/)
  - Custom-tools agent variant: Separate id gemini-3.1-pro-preview-customtools tuned for agentic workflows using custom tools and bash. (https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview)
  - Code-generated visuals: Showcased animated SVG generation, live dashboards and interactive 3D experiences from prompts. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/)

### Veo 3.1
**Veo 3.1** (Google DeepMind; preview; video-gen; released 2025-10-15) | Gemini API: `veo-3.1-generate-preview`; Gemini API: `veo-3.1-fast-generate-preview`; Gemini API: `veo-3.1-lite-generate-preview` | Web app: https://gemini.google.com — Specs: 4/6/8 s, 720p/1080p/4K (1080p/4K need 8 s; no 4K on Lite), 16:9 or 9:16, 24 fps. Standard + Fast released 2025-10-15, Lite 2026-03-31; all still preview ids in the Gemini API. Veo 2 and Veo 3.0 sunset 2026-06-30. Google now recommends Gemini Omni Flash as default video model.
  - Gemini API: `veo-3.1-generate-preview` — https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning (docs: https://ai.google.dev/gemini-api/docs/veo)
  - Gemini API: `veo-3.1-fast-generate-preview` — https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-fast-generate-preview:predictLongRunning
  - Gemini API: `veo-3.1-lite-generate-preview` — https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-lite-generate-preview:predictLongRunning
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/veo/3-1-generate)
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Native audio in every clip: Dialogue, SFX and ambience generated with the video; Veo 3.1 extended audio to Ingredients-to-Video, Frames-to-Video and Extend. (https://blog.google/innovation-and-ai/products/veo-updates-flow/)
  - Reference images and first/last frame control: Multiple reference images for character/object/style consistency; generate a bridge between a start and end frame. (https://blog.google/innovation-and-ai/products/veo-updates-flow/)
  - Video extension to a minute+: Extend clips (720p) to build longer continuous scenes. (https://ai.google.dev/gemini-api/docs/veo)

### Genie 3
**Genie 3** (Google DeepMind; preview; world-model; released 2025-08) | Project Genie (Google Labs): https://labs.google/projectgenie — No public API. Research preview announced Aug 2025; consumer access via Project Genie only for Google AI Ultra subscribers, US, 18+ (not Business accounts). Exact announcement day not re-verified.
  - Project Genie (Google Labs) — https://labs.google/projectgenie (docs: https://deepmind.google/models/genie/)
Notable capabilities:
  - [FIRST] Real-time interactive world generation: Generates navigable, photorealistic 720p worlds at 20-24 fps from text/image prompts. (https://deepmind.google/models/genie/)
  - World memory / consistency: Regions stay consistent when revisited; multi-minute visual consistency. (https://deepmind.google/models/genie/)
  - Promptable world events: Change weather or introduce objects/characters mid-exploration via text. (https://deepmind.google/models/genie/)
  - Consumer world sketching and remixing (found after launch): Project Genie (2026-01-29) lets users sketch, explore and remix worlds (60 s sessions), combining Genie 3 with Nano Banana Pro and Gemini. (https://blog.google/innovation-and-ai/models-and-research/google-deepmind/project-genie/)

### Lyria RealTime
**Lyria RealTime** (Google DeepMind; preview; music) | Gemini API (Live music, WebSocket): `models/lyria-realtime-exp` — Experimental model (status 'Experimental' on the Gemini API models page; no shutdown date announced). Instrumental only; output is SynthID-watermarked. No price listed on the Gemini API pricing page as of 2026-09-29. Release date not re-verified here (it first appeared in 2025 as an experimental model).
  - Gemini API (Live music, WebSocket): `models/lyria-realtime-exp` (docs: https://ai.google.dev/gemini-api/docs/realtime-music-generation)
Notable capabilities:
  - Interactive streaming music generation: Persistent bidirectional WebSocket session that continuously streams 48 kHz stereo 16-bit PCM; steer live with weighted text prompts and play/pause/stop/reset controls. (https://ai.google.dev/gemini-api/docs/realtime-music-generation)
  - Live musical parameters: Adjust guidance (0-6), BPM (60-200), density, brightness, scale (12 key pairs) and mute bass/drums on the fly. (https://ai.google.dev/gemini-api/docs/realtime-music-generation)

### Gemini 3.7 Flash
**Gemini 3.7 Flash** (Google DeepMind; legacy; reasoning-llm; released 2026-08-13) | ctx 1,048,576 | $0.75 in / $3.75 out per 1M tokens (Standard tier, introductory through 2026-12-31; $1.50 / $7.50 from 2027-01-01) | Gemini API: `gemini-3.7-flash`; OpenRouter: `google/gemini-3.7-flash` — Still served and stable (no shutdown date) but superseded by Gemini 3.8 Flash at the same price. Vertex model id not verified.
  - Gemini API: `gemini-3.7-flash` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-7-flash)
  - OpenRouter: `google/gemini-3.7-flash`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Production-quality coding: 43.6% FrontierCode 1.1 Main and 65.3% DeepSWE v1.1 (vs 34.4% / 49.0% for 3.6 Flash); WebDev Arena Elo 1588. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)
  - Enterprise document/automation work: 34.0% GDP.pdf and 30.4% AutomationBench, large jumps over 3.6 Flash. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/)
  - Half-price workhorse: Launched at half the original 3.6 Flash per-token price; thinking levels low/medium/high (minimal returns an error). (https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash)

### Gemini 3.6 Flash
**Gemini 3.6 Flash** (Google DeepMind; legacy; reasoning-llm; released 2026-07-21) | ctx 1,048,576 | $0.75 in / $3.75 out per 1M tokens (Standard tier, current introductory price through 2026-12-31; $1.50 / $7.50 from 2027-01-01, which was its launch price) | Gemini API: `gemini-3.6-flash`; OpenRouter: `google/gemini-3.6-flash` — Stable, no shutdown date; superseded by 3.7 and 3.8 Flash. Recommended replacement for gemini-3-flash-preview per deprecations page. Output limit not re-checked (assumed 65,536 like siblings, omitted).
  - Gemini API: `gemini-3.6-flash` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.6-flash:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-6-flash)
  - OpenRouter: `google/gemini-3.6-flash`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Token-efficient agentic coding: Uses 17% fewer output tokens than 3.5 Flash with better coding/multimodal results and fewer unwanted edits. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)
  - Computer-use agents: 83.0% on OSWorld-Verified (vs 78.4% for 3.5 Flash). (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)

### Gemini 3.5 Flash
**Gemini 3.5 Flash** (Google DeepMind; legacy; reasoning-llm; released 2026-05-19) | ctx 1,048,576 | $1.5 in / $9 out per 1M tokens (Standard tier) | Gemini API: `gemini-3.5-flash`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-3.5-flash`; OpenRouter: `google/gemini-3.5-flash` — Launched at Google I/O 2026 as first Gemini 3.5 model. Model page lists gemini-3-flash-preview (Dec 2025) as its preview predecessor id; that preview is still served. Now more expensive than 3.6-3.8 Flash; use gemini-3.8-flash.
  - Gemini API: `gemini-3.5-flash` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-3.5-flash` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-5-flash)
  - OpenRouter: `google/gemini-3.5-flash`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Flash beats previous Pro on agents: Outperformed Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo) and MCP Atlas (83.6%); 84.2% CharXiv Reasoning. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/)
  - High output speed: Google claims ~4x the output tokens/second of other frontier models at launch. (https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/)

### Gemini 3.1 Flash TTS (preview)
**Gemini 3.1 Flash TTS (preview)** (Google DeepMind; legacy; audio/speech; released 2026-04-15) | $1 in / $20 out per 1M tokens (USD), text in / audio out (25 audio tokens per second) | Gemini API: `gemini-3.1-flash-tts-preview`; Google Cloud Text-to-Speech (Gemini-TTS): `Gemini 3.1 Flash TTS (Preview)` — Superseded by gemini-3.8-flash-tts (GA 2026-09-22), which is cheaper at intro pricing ($0.50/$9.00). Still served as preview; also billed in Cloud TTS at the same $1/$20.
  - Gemini API: `gemini-3.1-flash-tts-preview` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-tts-preview:generateContent (docs: https://ai.google.dev/gemini-api/docs/models)
  - Google Cloud Text-to-Speech (Gemini-TTS): `Gemini 3.1 Flash TTS (Preview)` (docs: https://cloud.google.com/text-to-speech/pricing)
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Steerable expressive TTS: 'Cost-efficient, expressive, and steerable text to speech' controlled with natural-language prompts. (https://ai.google.dev/gemini-api/docs/changelog)

### Gemini 3.1 Flash Live (preview)
**Gemini 3.1 Flash Live (preview)** (Google DeepMind; legacy; audio/speech; released 2026-03-26) | $0.75 in / $4.5 out per 1M tokens (USD); same price as gemini-3.8-live | Gemini Live API (WebSocket): `gemini-3.1-flash-live-preview` — Preview id; the models page labels it legacy and recommends gemini-3.8-live (GA 2026-09-15). No shutdown date announced as of 2026-09-29. Context window not re-checked.
  - Gemini Live API (WebSocket): `gemini-3.1-flash-live-preview` (docs: https://ai.google.dev/gemini-api/docs/models)
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Audio-to-audio real-time dialogue: Native audio model 'designed for real-time dialogue and voice-first AI applications' on the Live API. (https://ai.google.dev/gemini-api/docs/changelog)

### Lyria 3 (Clip / Pro)
**Lyria 3 (Clip / Pro)** (Google DeepMind; legacy; music; released 2026-02-18) | Gemini API (Interactions API): `lyria-3-clip-preview`; Gemini API (Interactions API): `lyria-3-pro-preview`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-3-pro-preview`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-3-clip-preview`; OpenRouter: `google/lyria-3-pro-preview` | Web app: https://gemini.google.com — Lyria 3 launched 2026-02-18 in the Gemini app (30 s clips) and YouTube Dream Track; Lyria 3 Pro and the developer previews (lyria-3-clip-preview, lyria-3-pro-preview) followed on 2026-03-25 (Gemini API, AI Studio, Vertex public preview, Google Vids, ProducerAI). Superseded by Lyria 3.5 (lyria-3.5, GA 2026-09-03); Gemini API pricing page now lists both as 'Lyria 3 legacy models'; no shutdown date announced. Artist names in prompts are treated as broad inspiration only.
  - Gemini API (Interactions API): `lyria-3-clip-preview` — https://generativelanguage.googleapis.com/v1beta/interactions (docs: https://ai.google.dev/gemini-api/docs/music-generation)
  - Gemini API (Interactions API): `lyria-3-pro-preview` — https://generativelanguage.googleapis.com/v1beta/interactions
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-3-pro-preview` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-3-clip-preview`
  - OpenRouter: `google/lyria-3-pro-preview`
  - Web app — https://gemini.google.com
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Songs with vocals and auto-written lyrics in the Gemini app: 30-second tracks with vocals and lyrics from a text prompt, photo or video, with Nano Banana cover art; 8 languages (en, de, es, fr, hi, ja, ko, pt); 18+ only. (https://blog.google/innovation-and-ai/products/gemini-app/lyria-3/)
  - SynthID watermark + detection in Gemini: All outputs carry SynthID; the Gemini app can check whether uploaded audio was generated with Google AI via SynthID. (https://blog.google/innovation-and-ai/products/gemini-app/lyria-3/)
  - Full songs with structure control (Lyria 3 Pro): Tracks up to ~3 minutes (184 s max on Vertex) with control over intros, verses, choruses and bridges, duration, BPM and intensity; 44.1 kHz, 192 kbps MP3; C2PA content credentials and vocal-likeness filtering. (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3)

### Lyria 2
**Lyria 2** (Google DeepMind; legacy; music; released 2025-10-27) | Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-002` — Vertex page lists lyria-002 as GA with release date 2025-10-27 (Lyria 2 was first shown publicly in 2025; earlier preview dates not re-verified). No vocals, lyrics or image input; superseded by Lyria 3 / 3.5 but still GA on Vertex, global region only. Status 'legacy' is our judgement (no deprecation announced).
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `lyria-002` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-002)
  - Pricing source: https://cloud.google.com/vertex-ai/generative-ai/pricing
Notable capabilities:
  - Instrumental clips with negative prompting: Text-to-music instrumental clips up to 32.8 s, 48 kHz WAV, up to 4 clips per prompt, negative prompts supported; US English prompts only. (https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-002)

### Gemini 2.5 Flash Native Audio (Live, preview)
**Gemini 2.5 Flash Native Audio (Live, preview)** (Google DeepMind; legacy; audio/speech; released 2025-09-23) | $0.5 in / $2 out per 1M tokens (USD); audio/video in $3.00 | Gemini Live API (WebSocket): `gemini-2.5-flash-native-audio-preview-12-2025`; Gemini Live API (WebSocket): `gemini-2.5-flash-native-audio-preview-09-2025` — Preview snapshots 2025-09-23 and 2025-12-12. No shutdown date announced; migrate to gemini-3.8-live. Older gemini-2.0-flash-live-001 and gemini-live-2.5-flash-preview were shut down 2025-12-09.
  - Gemini Live API (WebSocket): `gemini-2.5-flash-native-audio-preview-12-2025` (docs: https://ai.google.dev/gemini-api/docs/models)
  - Gemini Live API (WebSocket): `gemini-2.5-flash-native-audio-preview-09-2025` (docs: https://ai.google.dev/gemini-api/docs/changelog)
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Native-audio reasoning in the Live API: Low-latency voice and video agents with native audio reasoning; 09-2025 snapshot improved function calling and speech cut-off handling, 12-2025 snapshot improved complex workflows. (https://ai.google.dev/gemini-api/docs/changelog)

### Gemini 2.5 Flash-Lite
**Gemini 2.5 Flash-Lite** (Google DeepMind; legacy; llm; released 2025-07-22) | ctx 1,048,576 | $0.1 in / $0.4 out per 1M tokens (Standard; text/image/video input; audio input $0.30) | Gemini API: `gemini-2.5-flash-lite`; OpenRouter: `google/gemini-2.5-flash-lite` — GA 2025-07-22; no shutdown date, access limited to historical users; replacement 3.5 Flash-Lite. Max output not re-verified.
  - Gemini API: `gemini-2.5-flash-lite` — https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent (docs: https://ai.google.dev/gemini-api/docs/models)
  - OpenRouter: `google/gemini-2.5-flash-lite`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Cheapest Gemini text tier: Still the lowest per-token Gemini text price ($0.10 / $0.40) with 1M context. (https://ai.google.dev/gemini-api/docs/pricing)
  - Thinking off by default: Lowest latency/cost in the 2.5 family, thinking disabled by default, yet supports grounding, code execution, URL context and function calling. (https://developers.googleblog.com/en/gemini-2-5-thinking-model-updates/)

### Gemini 2.5 Flash
**Gemini 2.5 Flash** (Google DeepMind; legacy; reasoning-llm; released 2025-06-17) | ctx 1,048,576 | $0.3 in / $2.5 out per 1M tokens (Standard; text/image/video input; audio input $1.00) | Gemini API: `gemini-2.5-flash`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-2.5-flash`; OpenRouter: `google/gemini-2.5-flash` — Preview 2025-04-17, GA 2025-06-17. No shutdown date, but access limited to prior users; replacement 3.5 Flash-Lite or 3.8 Flash.
  - Gemini API: `gemini-2.5-flash` — https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-2.5-flash` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-flash)
  - OpenRouter: `google/gemini-2.5-flash`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Hybrid reasoning with thinking budget: Thinking can be controlled per request; 1M-token multimodal context at low price. (https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash)
  - First fully hybrid reasoning model (Google): Google's first model where thinking can be switched on/off, with a 0-24,576 token thinking budget. (https://developers.googleblog.com/en/start-building-with-gemini-25-flash/)

### Gemini 2.5 Pro
**Gemini 2.5 Pro** (Google DeepMind; legacy; reasoning-llm; released 2025-06-17) | ctx 1,048,576 | $1.25 in / $10 out per 1M tokens (Standard, prompts <=200k; >200k: $2.50 in / $15.00 out) | Gemini API: `gemini-2.5-pro`; Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-2.5-pro`; OpenRouter: `google/gemini-2.5-pro` — First released as experimental 2025-03; GA 2025-06-17 (stable id; earlier preview ids e.g. gemini-2.5-pro-preview-*). No shutdown date, but Gemini API access is limited to projects that used it before; Google recommends 3.5 Flash-Lite or 3.8 Flash for new projects.
  - Gemini API: `gemini-2.5-pro` — https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-pro:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI): `gemini-2.5-pro` (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-pro)
  - OpenRouter: `google/gemini-2.5-pro`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Thinking model with 1M context: Built-in thinking plus 1,048,576-token multimodal input and 65K output; Search/Maps grounding, code execution, URL context. (https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro)
  - Debuted: First Gemini 2.5 'thinking model'; the March 2025 experimental release topped LMArena by a significant margin and led coding/math/science benchmarks. (https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-model-thinking-updates-march-2025/)
  - Coding-agent backbone (found after launch): Steepest demand growth of any Google model; powered tools such as Cursor and GitHub Copilot at GA. (https://developers.googleblog.com/en/gemini-2-5-thinking-model-updates/)

### Gemini 2.5 Flash TTS / Pro TTS
**Gemini 2.5 Flash TTS / Pro TTS** (Google DeepMind; legacy; audio/speech; released 2025-05-20) | $0.5 in / $10 out per 1M tokens (USD) for Flash TTS; Pro TTS $1.00 in / $20.00 audio out (25 audio tokens per second) | Gemini API: `gemini-2.5-flash-preview-tts`; Gemini API: `gemini-2.5-pro-preview-tts`; Google Cloud Text-to-Speech (Gemini-TTS, GA): `gemini-2.5-flash-tts`; Google Cloud Text-to-Speech (Gemini-TTS, GA): `gemini-2.5-pro-tts`; Google Cloud Text-to-Speech (preview): `gemini-2.5-flash-lite-preview-tts` — Gemini API ids are 'Limited Access' preview with no shutdown date (migrate to gemini-3.8-flash-tts / -lite-tts). In Cloud TTS, gemini-2.5-flash-tts and gemini-2.5-pro-tts went GA 2025-09-30; streaming added 2025-11-07. Released date = Google I/O 2025 preview (from memory, not re-verified today); Dec 10 2025 update improved expressivity and pacing.
  - Gemini API: `gemini-2.5-flash-preview-tts` (docs: https://ai.google.dev/gemini-api/docs/models)
  - Gemini API: `gemini-2.5-pro-preview-tts` (docs: https://ai.google.dev/gemini-api/docs/models)
  - Google Cloud Text-to-Speech (Gemini-TTS, GA): `gemini-2.5-flash-tts` (docs: https://docs.cloud.google.com/text-to-speech/docs/release-notes)
  - Google Cloud Text-to-Speech (Gemini-TTS, GA): `gemini-2.5-pro-tts`
  - Google Cloud Text-to-Speech (preview): `gemini-2.5-flash-lite-preview-tts`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Prompt-controlled multi-speaker TTS: Natural-language control of style, accent, pace and emotion; single and multi-speaker synthesis; 30 speakers in 80+ locales (Cloud GA). (https://docs.cloud.google.com/text-to-speech/docs/release-notes)

### Gemini 3.1 Flash-Lite
**Gemini 3.1 Flash-Lite** (Google DeepMind; deprecated; llm; released 2026-05-07) | ctx 1,048,576 | $0.25 in / $1.5 out per 1M tokens (Standard; text/image/video input; audio input $0.50) | Gemini API: `gemini-3.1-flash-lite`; OpenRouter: `google/gemini-3.1-flash-lite` — Stable GA 2026-05-07; scheduled shutdown 2027-05-07, replacement gemini-3.5-flash-lite. Preview id gemini-3.1-flash-lite-preview (early 2026) still listed by the live API / OpenRouter though docs list it as shut down.
  - Gemini API: `gemini-3.1-flash-lite` — https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite:generateContent (docs: https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite)
  - OpenRouter: `google/gemini-3.1-flash-lite`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Low-cost frontier-class Lite: Described as frontier-class performance at reduced cost; cheapest per-token 3.x text model. (https://ai.google.dev/gemini-api/docs/models)
  - Full tool stack on a Lite model: 1M-token multimodal input (text, image, video, audio, PDF) with 65K output. (https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite)

### Nano Banana (Gemini 2.5 Flash Image)
**Nano Banana (Gemini 2.5 Flash Image)** (Google DeepMind; deprecated; image-gen; released 2025-10-02) | $0.3 in / $? out input per 1M tokens; image output $0.039 per image | Gemini API: `gemini-2.5-flash-image`; OpenRouter: `google/gemini-2.5-flash-image` — Stable GA 2025-10-02 (preview 2025-08-26); SHUTS DOWN 2026-10-02, replacement gemini-3.1-flash-image.
  - Gemini API: `gemini-2.5-flash-image` — https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:generateContent (docs: https://ai.google.dev/gemini-api/docs/models)
  - Google Cloud Gemini Enterprise Agent Platform (Vertex AI) (docs: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-flash-image)
  - OpenRouter: `google/gemini-2.5-flash-image`
  - Pricing source: https://ai.google.dev/gemini-api/docs/pricing
Notable capabilities:
  - Conversational image editing: The original 'Nano Banana': multi-turn natural-language image editing with character consistency, which made Gemini image editing go viral in 2025. (https://ai.google.dev/gemini-api/docs/models)
  - Multi-image fusion and targeted edits: Blend multiple images, keep characters consistent, and do prompt-based local edits (background blur, object removal, colorization); SynthID on all outputs. (https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/)

### Gemini Robotics-ER 1.5 / 1.6
**Gemini Robotics-ER 1.5 / 1.6** (Google DeepMind; retired; robotics; released 2025-09-25) | ctx 131,072 | Gemini API (shut down): `gemini-robotics-er-1.6-preview`; Gemini API (shut down): `gemini-robotics-er-1.5-preview` — Retired. gemini-robotics-er-1.5-preview released 2025-09-25, shut down 2026-04-30 (replaced by 1.6). gemini-robotics-er-1.6-preview released 2026-04-14, shut down 2026-08-31 (replaced by gemini-robotics-er-2-preview). Token limits and Jan 2025 cutoff are those listed for ER 1.6 on the Gemini API model page.
  - Gemini API (shut down): `gemini-robotics-er-1.6-preview` (docs: https://ai.google.dev/gemini-api/docs/deprecations)
  - Gemini API (shut down): `gemini-robotics-er-1.5-preview` (docs: https://ai.google.dev/gemini-api/docs/deprecations)
Notable capabilities:
  - Embodied reasoning in the public Gemini API: ER 1.5 (2025-09-25) exposed embodied reasoning (pointing, 2D boxes, trajectories, task planning, tool calls) available in the public Gemini API, while the VLA stayed partner-only. (https://ai.google.dev/gemini-api/docs/deprecations)
  - Instrument reading (ER 1.6): Reads pressure gauges, thermometers, sight glasses and digital readouts: 86% (93% with agentic vision) vs 23% for ER 1.5 and 67% for Gemini 3 Flash; built with Boston Dynamics and used by Spot for inspections. (https://deepmind.google/blog/gemini-robotics-er-1-6/)

### Imagen 4
**Imagen 4** (Google DeepMind; retired; image-gen; released 2025-06-24) | Gemini API (shut down): `imagen-4.0-generate-001` — Imagen 4.0 variants released 2025-06-24, shut down in the Gemini API 2026-08-17; replacement gemini-3.1-flash-image. Other variant ids (fast/ultra) and Vertex status not verified; pricing not verified (retired).
  - Gemini API (shut down): `imagen-4.0-generate-001` (docs: https://ai.google.dev/gemini-api/docs/deprecations)
Notable capabilities:
  - Dedicated text-to-image diffusion model: Google's last standalone Imagen generation; superseded by Gemini-native image models (Nano Banana 2). (https://ai.google.dev/gemini-api/docs/deprecations)
  - Retired in favor of Gemini-native imaging (found after launch): Deprecation table names gemini-3.1-flash-image as replacement, marking the shift from standalone diffusion models to Gemini image models. (https://ai.google.dev/gemini-api/docs/deprecations)


## hexgrad

### Kokoro-82M
**Kokoro-82M** (hexgrad; current; audio/speech; released 2025-01-27; open weights) | Hugging Face: `hexgrad/Kokoro-82M`; pip: `kokoro`; DeepInfra: `hexgrad/Kokoro-82M` | OpenRouter: https://openrouter.ai/hexgrad/kokoro-82m — v1.0: 54 preset voices, 8 languages (US/UK English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, Mandarin), 24 kHz. No voice cloning. v0.19 was 2024-12-25. ~11.5M HF downloads/month.
  - Hugging Face: `hexgrad/Kokoro-82M` — https://huggingface.co/hexgrad/Kokoro-82M
  - pip: `kokoro` — https://github.com/hexgrad/kokoro
  - DeepInfra: `hexgrad/Kokoro-82M` — https://deepinfra.com/hexgrad/Kokoro-82M
  - OpenRouter — https://openrouter.ai/hexgrad/kokoro-82m
  - Pricing source: https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice
Notable capabilities:
  - Tiny model, top-tier quality: 82M-param StyleTTS2 + ISTFTNet model trained for ~$1,000 (1,000 A100 h) on permissive data; v0.19 hit #1 on the HF TTS Spaces Arena; still top-5 open weights on Artificial Analysis (~1065 Elo) in Sept 2026. (https://huggingface.co/hexgrad/Kokoro-82M)


## Hugging Face

### SmolVLA (450M)
**SmolVLA (450M)** (Hugging Face; current; robotics; released 2025-06-03; open weights) | Hugging Face: `lerobot/smolvla_base` | GitHub (LeRobot): https://github.com/huggingface/lerobot — Designed for low-cost arms (SO-100/SO-101). HF repo still updated in Sept 2026; variants lerobot/smolvla_libero, lerobot/smolvla_robotwin. NVIDIA announced it would acquire Hugging Face (see 2026-09-03 entry).
  - Hugging Face: `lerobot/smolvla_base` — https://huggingface.co/lerobot/smolvla_base (docs: https://huggingface.co/docs/lerobot/smolvla)
  - GitHub (LeRobot) — https://github.com/huggingface/lerobot
Notable capabilities:
  - VLA small enough for a laptop: 450M params (SmolVLM2-500M backbone + flow-matching action expert); trains on a single GPU and runs on consumer hardware incl. MacBooks. (https://huggingface.co/blog/smolvla)
  - Trained on community-shared data: Pretrained on ~10M frames from 487 community LeRobot datasets (<30k episodes, an order of magnitude less than other VLAs); 78.3% success on real SO-100 tasks. (https://huggingface.co/blog/smolvla)
  - Asynchronous inference: Decouples action prediction from execution: ~30% faster task completion and 2x throughput. (https://huggingface.co/blog/smolvla)


## Hume AI

### Hume Octave 2 (TTS)
**Hume Octave 2 (TTS)** (Hume AI; current; audio/speech; released 2025-10-01) | Hume API: `version: 2` | Web app: https://platform.hume.ai — Select via `version: 2` in the TTS request body (`1` = Octave 1, English/Spanish, ~200 ms). Docs still label Octave 2 '(preview)' as of 2026-09-29. Auth header X-Hume-Api-Key. Max 5,000 chars per utterance, 1,000-char descriptions. Formats MP3/WAV/PCM. No Octave 3 announced on Hume's blog through Sept 2026. Speech-to-speech sibling: see hume-evi.
  - Hume API: `version: 2` — https://api.hume.ai/v0/tts (docs: https://dev.hume.ai/docs/text-to-speech-tts/overview)
  - Web app — https://platform.hume.ai
  - Pricing source: https://www.hume.ai/pricing
Notable capabilities:
  - LLM-based emotionally intelligent TTS: Speech-language model that infers emotion and delivery from text; natural-language 'acting instructions' steer tone. Octave 2 at half the price of Octave 1, ~100 ms model latency (docs) / under 200 ms (launch blog). (https://www.hume.ai/blog/octave-2-launch)
  - 11 languages, instant cloning from 15 s: Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian, Spanish; instant voice cloning from a ~15 s recording with accent prediction across languages; voice design from a text prompt (English only). (https://dev.hume.ai/docs/text-to-speech-tts/overview)
  - Voice conversion and phoneme editing: Launch post describes voice conversion (swap speaker) and direct phoneme-level pronunciation editing as new capabilities for a speech-language model. (https://www.hume.ai/blog/octave-2-launch)

### Hume EVI 3 / EVI 4 mini (speech-to-speech)
**Hume EVI 3 / EVI 4 mini (speech-to-speech)** (Hume AI; current; audio/speech; released 2025-05-29) | Hume API (EVI WebSocket): `EVI version 3 or 4-mini (set in EVI config)` | Web app: https://platform.hume.ai — EVI 3 is English-only and can answer without an external LLM ('quick responses'); EVI 4 mini is multilingual but requires a supplemental LLM. Both share the same WebSocket; version is chosen in the EVI configuration. Full EVI 4 not launched as of 2026-09-29 (not on Hume blog). EVI 1/2 are older generations.
  - Hume API (EVI WebSocket): `EVI version 3 or 4-mini (set in EVI config)` — wss://api.hume.ai/v0/evi/chat (docs: https://dev.hume.ai/docs/speech-to-speech-evi/overview)
  - Web app — https://platform.hume.ai
  - Pricing source: https://www.hume.ai/pricing
Notable capabilities:
  - Empathic voice interface with any prompted voice: EVI 3 (2025-05-29) is a speech-to-speech foundation model that can speak in any of 100,000+ custom voices created via prompting, with inferred personality; ~1.2 s practical end-of-speech-to-response latency at launch. (https://www.hume.ai/blog/introducing-evi-3)
  - EVI 4 mini: Octave 2 voice in 11 languages: EVI 4 mini (announced with Octave 2 on 2025-10-01) brings Octave 2 to the speech-to-speech API in 11 languages but must be paired with an external LLM (Anthropic, OpenAI, Google, Fireworks...) until the full EVI 4 ships. (https://www.hume.ai/blog/octave-2-launch)


## Ideogram

### Ideogram 4.0
**Ideogram 4.0** (Ideogram; current; image-gen; released 2026-06-03; open weights) | Ideogram API: `ideogram-v4` | Hugging Face: https://huggingface.co/ideogram-ai/ideogram-4-fp8; Hugging Face (NF4): https://huggingface.co/ideogram-ai/ideogram-4-nf4; Web app: https://ideogram.ai — Some third-party sites claim Apache-2.0 - HF card says license: other (non-commercial). API: Api-Key header, multipart with text_prompt or json_prompt; rendering_speed=FLASH currently returns 400. Also /v1/ideogram-v3/generate (previous gen). Per-image API pricing not verified on official page.
  - Ideogram API: `ideogram-v4` — https://api.ideogram.ai/v1/ideogram-v4/generate (docs: https://developer.ideogram.ai/api-reference)
  - Hugging Face — https://huggingface.co/ideogram-ai/ideogram-4-fp8
  - Hugging Face (NF4) — https://huggingface.co/ideogram-ai/ideogram-4-nf4
  - Web app — https://ideogram.ai
Notable capabilities:
  - Structured JSON prompting with layout control: Native JSON prompt format with explicit bounding-box layout and color-palette controls. (https://ideogram.ai/blog/ideogram-4.0/)
  - Best-in-class multilingual text rendering: Strong in-image typography across languages; native 2K resolution. (https://ideogram.ai/blog/ideogram-4.0/)
  - First Ideogram open-weight model: 9.3B DiT trained from scratch, Qwen3-VL-8B text encoder; quantized weights on HF for research. (https://huggingface.co/ideogram-ai/ideogram-4-fp8)


## Inworld AI

### Inworld Realtime TTS-2 / TTS-2 Flash
**Inworld Realtime TTS-2 / TTS-2 Flash** (Inworld AI; current; audio/speech; released 2026-08-31) | Inworld API: `inworld-tts-2` | Cloudflare Workers AI: https://developers.cloudflare.com/ai/models/inworld/tts-2/ — Research preview 2026-05-05, GA 2026-08-31. Inworld claimed #1 on Artificial Analysis Speech Arena; on 2026-09-29 AA shows it #5 (Elo 1244) behind Eleven v4, Sonic 3.6, Gemini 3.8 Flash TTS, Qwen-Audio-3.0-TTS-Plus. Docs say 200+ languages vs 100+ in blog. TTS-1..1.5 discontinued 2026-06-15 (auto-routed). Flash model id not verified. Max 2,000 chars/request.
  - Inworld API: `inworld-tts-2` — https://api.inworld.ai/tts/v1/voice (docs: https://docs.inworld.ai/tts/tts-models)
  - Cloudflare Workers AI — https://developers.cloudflare.com/ai/models/inworld/tts-2/
  - Pricing source: https://inworld.ai/pricing
Notable capabilities:
  - Closed-loop, audio-aware delivery: Conditions on the actual audio of prior turns (user tone, pacing, emotion), not just transcripts, and takes plain-English voice direction; delivery modes STABLE/BALANCED/CREATIVE. (https://inworld.ai/blog/realtime-tts-2)
  - Cross-lingual identity in 100+ languages: One voice holds identity while switching language on the fly; cloning from 5-15 s reference or voice design from a text description. (https://inworld.ai/blog/realtime-tts-2)
  - Flash variant ~20 ms TTFB: TTS-2 Flash: ~20 ms TTFB, ~5x faster than inworld-tts-2 (docs); TTS-2 median TTFA under 200 ms. (https://docs.inworld.ai/tts/tts-models)


## Kuaishou

### Kling 3.0 (VIDEO 3.0 / 3.0 Omni)
**Kling 3.0 (VIDEO 3.0 / 3.0 Omni)** (Kuaishou; current; video-gen; released 2026-02) | fal.ai: `fal-ai/kling-video/v3/standard/text-to-video` | Web app: https://kling.ai — Released early Feb 2026 (official guide says Feb 6; other sources Feb 7). Official API model_name strings not verified (docs are JS-rendered); fal ids verified: kling-video/v3/{standard,pro}/{text,image}-to-video, plus turbo/4K variants. Pricing not verified.
  - Kling AI API — https://api-singapore.klingai.com (docs: https://kling.ai/document-api/quickStart/productIntroduction/overview)
  - fal.ai: `fal-ai/kling-video/v3/standard/text-to-video` (docs: https://fal.ai/models/fal-ai/kling-video/v3/standard/text-to-video/api)
  - Web app — https://kling.ai
Notable capabilities:
  - Multi-shot storyboards: Generates multi-shot narrative sequences in one job, with storyboard control over shots. (https://kling.ai/quickstart/klingai-video-3-model-user-guide)
  - Native multilingual audio: Native audio (dialogue/SFX) generated with the video, multilingual; clips up to 15 s. (https://kling.ai/quickstart/klingai-video-3-model-user-guide)
  - Unified Omni model with element consistency: VIDEO 3.0 Omni (successor of O1) unifies generation and editing with stronger element/character consistency; IMAGE 3.0 / 3.0 Omni siblings. (https://kling.ai/quickstart/klingai-video-3-model-user-guide)


## Kunlun Tech (Skywork AI)

### Mureka V9.5 (and O3)
**Mureka V9.5 (and O3)** (Kunlun Tech (Skywork AI); current; music; released 2026-07) | Mureka API: `mureka-9.5` | Web app: https://www.mureka.ai — Release history per the API changelog (https://platform.mureka.ai/docs/en/changelog.html): mureka-7 + mureka-o1 2025-07-29; mureka-7.5 2025-09-25; mureka-7.6 + mureka-o2 2025-12-09; mureka-8 2026-03-02 (consumer Mureka V8 announced 2026-01-28, claimed to surpass Suno in melody, vocals, arrangement and emotion; cited as a baseline in Tencent's SongGeneration 2 paper); mureka-9 2026-04-09; enhanced mureka-9.5 2026-08-28. V9.5 was shown around WAIC (late July 2026) and formally announced 2026-08-31 (GlobeNewswire) with internal-test figures: 61.0% of lead vocals rated convincing, 97.0% prompt following, 95.7% genre match. Exact consumer launch day and API pricing not verified; training-data provenance undisclosed. Kunlun Tech's music models are developed under its Skywork AI unit.
  - Mureka API: `mureka-9.5` (docs: https://platform.mureka.ai/docs/)
  - Web app — https://www.mureka.ai
Notable capabilities:
  - MusiCoT (music chain-of-thought) planning: Mureka's line plans song structure, sections and intent before generating audio (MusiCoT); Mureka O1 (2025-07-29) was billed as the first 'thinking' music reasoning model, followed by O2 (2025-12-09) and O3 'reflective reasoning' with V9.5. (https://www.prnewswire.com/news-releases/kunlun-tech-launches-the-worlds-first-music-reasoning-large-model-mureka-o1-leading-the-global-ai-music-revolution-302411665.html)
  - MuCo creation agent: Agent that manages a song as a version-controlled project instead of one-shot generation (per Pandaily/Variety coverage of V9.5). (https://pandaily.com/mureka-v9-5-ai-music-kunlun-tech-jul2026)
  - Fine-tuning API and vocal cloning: API offers song/instrumental/lyrics generation, song extension, stem separation, transcription, vocal cloning and custom-model fine-tuning on 200+ consistent tracks. (https://platform.mureka.ai/docs/)


## Kyutai

### Kyutai Pocket TTS
**Kyutai Pocket TTS** (Kyutai; current; audio/speech; released 2026-01-13; open weights) | Hugging Face: `kyutai/pocket-tts`; Hugging Face (no cloning variant): `kyutai/pocket-tts-without-voice-cloning`; GitHub / pip: `pocket-tts` — Gated on HF (accept prohibited-use terms). Training code released 2026-08-25; 2026-09-28 post describes a 'drifting' objective replacing flow matching for the sampler head. `pip install pocket-tts`. Community WebAssembly ports run in-browser.
  - Hugging Face: `kyutai/pocket-tts` — https://huggingface.co/kyutai/pocket-tts
  - Hugging Face (no cloning variant): `kyutai/pocket-tts-without-voice-cloning` — https://huggingface.co/kyutai
  - GitHub / pip: `pocket-tts` — https://github.com/kyutai-labs/pocket-tts
Notable capabilities:
  - 100M-param TTS with cloning, real time on CPU: ~200 ms to first audio and ~6x real time on a MacBook Air M4 CPU; streaming, unbounded text length; voice cloning from audio. (https://huggingface.co/kyutai/pocket-tts)
  - Six languages (found after launch): English, French, German, Spanish, Portuguese, Italian (multilingual since 2026-05-04). (https://kyutai.org/blog/)

### Kyutai TTS 1.6B / Kyutai STT + Unmute
**Kyutai TTS 1.6B / Kyutai STT + Unmute** (Kyutai; current; audio/speech; released 2025-07-03; open weights) | Hugging Face (TTS): `kyutai/tts-1.6b-en_fr`; Hugging Face (STT): `kyutai/stt-2.6b-en`; Hugging Face (STT): `kyutai/stt-1b-en_fr` | GitHub (Unmute): https://github.com/kyutai-labs/unmute — STT open-sourced 2025-06-19, TTS + Unmute open-sourced 2025-07-03 (Kyutai blog). Weights CC-BY-4.0. For CPU TTS see kyutai-pocket-tts.
  - Hugging Face (TTS): `kyutai/tts-1.6b-en_fr` — https://huggingface.co/kyutai/tts-1.6b-en_fr
  - Hugging Face (STT): `kyutai/stt-2.6b-en` — https://huggingface.co/kyutai/stt-2.6b-en
  - Hugging Face (STT): `kyutai/stt-1b-en_fr` — https://huggingface.co/kyutai/stt-1b-en_fr
  - GitHub (Unmute) — https://github.com/kyutai-labs/unmute
Notable capabilities:
  - Text-streaming TTS: Delayed-streams architecture (~1.8B params incl. 600M depth transformer) starts speaking before the full text is available, English + French; voices only via pre-computed embeddings (no raw cloning, by design). (https://huggingface.co/kyutai/tts-1.6b-en_fr)
  - Streaming STT with semantic VAD: stt-2.6b-en (English, 2.5 s delay) and stt-1b-en_fr (0.5 s delay) transcribe as audio arrives; used in Unmute, which wraps any text LLM with real-time STT+TTS. (https://huggingface.co/kyutai/stt-2.6b-en)

### Kyutai Moshi / Hibiki-Zero (full-duplex speech models)
**Kyutai Moshi / Hibiki-Zero (full-duplex speech models)** (Kyutai; current; audio/speech; released 2024-09-17; open weights) | Hugging Face: `kyutai/moshiko-pytorch-bf16`; Hugging Face: `kyutai/hibiki-zero-3b-pytorch-bf16` | Web demo: https://moshi.chat — Moshi (announced July 2024, weights + paper Sept 2024) is widely cited as the first real-time full-duplex open spoken dialogue model; NVIDIA PersonaPlex-7B (Jan 2026) is fine-tuned from Moshiko weights. Variants: moshiko (male)/moshika (female) in PyTorch bf16/int8, MLX int4/int8/bf16, Rust/Candle. Code MIT/Apache, weights CC-BY-4.0.
  - Hugging Face: `kyutai/moshiko-pytorch-bf16` — https://github.com/kyutai-labs/moshi
  - Hugging Face: `kyutai/hibiki-zero-3b-pytorch-bf16` — https://huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16
  - Web demo — https://moshi.chat
Notable capabilities:
  - [FIRST] Open full-duplex spoken dialogue: 7B temporal transformer modelling user and Moshi audio streams simultaneously with an 'inner monologue' text stream; 160 ms theoretical / ~200 ms practical latency on an L4; Mimi codec (24 kHz, 12.5 Hz, 1.1 kbps). (https://github.com/kyutai-labs/moshi)
  - Hibiki-Zero simultaneous speech translation (found after launch): 3B model (2026-02-12) translating French, Spanish, Portuguese and German speech to English in real time with voice transfer, trained without aligned data. (https://kyutai.org/blog/)
  - MoshiRAG (found after launch): Asynchronous knowledge retrieval via a text LLM for full-duplex speech models (2026-04-30); RL post-training for interactivity (2026-06-10). (https://kyutai.org/blog/)


## Luma AI

### Luma Ray3.2
**Luma Ray3.2** (Luma AI; current; video-gen; released 2026-06-09) | Luma API: `ray-3.2` | Web app (Dream Machine): https://app.lumalabs.ai — Successor of Ray3 / Ray3 Modify / Ray3.14. Same API also serves image models uni-1 and uni-1-max (UNI-1.1). Credit-based API pricing (https://lumalabs.ai/pricing) - per-second price not verified.
  - Luma API: `ray-3.2` — https://agents.lumalabs.ai/v1/generations (docs: https://docs.agents.lumalabs.ai/)
  - Web app (Dream Machine) — https://app.lumalabs.ai
Notable capabilities:
  - Multi-keyframe direction: Up to 16 keyframes inside a single clip for frame-level control of how action evolves. (https://lumalabs.ai/news/introducing-ray-3-2)
  - Native HDR with 16-bit EXR export: Generates native HDR video with 16-bit EXR export for pro post-production; up to 20 s at 1080p. (https://lumalabs.ai/news/introducing-ray-3-2)
  - Multi-face performance tracking and reframe: Performance tracking for up to 8 faces and an improved reframe tool; full Ray control surface exposed via API for the first time. (https://lumalabs.ai/news/introducing-ray-3-2)


## Meta

### Muse Voice Transcribe 1.0
**Muse Voice Transcribe 1.0** (Meta; current; audio/speech; released 2026-09-03) | Meta Model API (streaming): `muse-voice-transcribe-1.0`; Meta Model API (file): `muse-voice-transcribe-1.0` — Meta's first real-time audio perception model on the Meta Model API (launched 2026-09-03); 25+ languages. Speech-to-text only: Meta does not offer a TTS or speech-to-speech API; Muse's realtime voice mode and Muse Realtime Avatar (Connect, 2026-09-23) are consumer features without a documented API.
  - Meta Model API (streaming): `muse-voice-transcribe-1.0` — wss://api.meta.ai/v1/asr/realtime (docs: https://dev.meta.ai/docs/overview)
  - Meta Model API (file): `muse-voice-transcribe-1.0` — https://api.meta.ai/v1/asr/transcribe (docs: https://dev.meta.ai/docs/overview)
  - Pricing source: https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/
Notable capabilities:
  - #1 streaming STT on Artificial Analysis (claimed): Meta says it ranks first on the Artificial Analysis streaming speech-to-text leaderboard and had the lowest average diarization error rate among APIs tested, streaming and offline. (https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/)
  - Diarization, VAD and endpointing in one model: Speaker attribution for 20+ speakers, punctuation, speech-boundary detection and adaptive delay (uses more audio context only for ambiguous words). (https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/)

### Muse Spark 1.3
**Muse Spark 1.3** (Meta; current; reasoning-llm; released 2026-09-02) | ctx 1,000,000 | $1.25 in / $4.25 out per 1M tokens (USD), standard tier; "contributor" tier muse-spark-1.3-contributor is $0.10/$0.002 cached/$0.20 | Meta Model API: `muse-spark-1.3`; OpenRouter: `meta/muse-spark-1.3` | Web app: https://meta.ai — Other ids: muse-spark-1.2, muse-spark-1.1, muse-spark-1.3-contributor, muse-spark-1.2-contributor. OpenAI-SDK-compatible API (public preview, self-serve). Contributor tier data terms not verified. Knowledge cutoff not published.
  - Meta Model API: `muse-spark-1.3` — https://api.meta.ai/v1/chat/completions (docs: https://dev.meta.ai/docs/)
  - OpenRouter: `meta/muse-spark-1.3` — https://openrouter.ai/meta/muse-spark-1.3
  - Web app — https://meta.ai
  - Pricing source: https://dev.meta.ai/models/muse-spark/
Notable capabilities:
  - Closed-weights successor to Llama: Proprietary model from Meta Superintelligence Labs; Muse Spark replaced Llama in Meta AI in April 2026. (https://venturebeat.com/technology/goodbye-llama-meta-launches-new-proprietary-ai-model-muse-spark-first-since)
  - Native video + document perception: Natively multimodal input (video, images, documents, text) with 1M context and 200K max output. (https://dev.meta.ai/models/muse-spark/)
  - Long-horizon multi-agent tuning: 1.3 tuned for long-running, multi-agent agentic builds; also powers Meta's Muse Code. (https://x.com/MetaforDevs/status/2095232442953236714)
  - Contributor pricing tier: Separate -contributor model ids priced ~90% lower (data-sharing tier). (https://dev.meta.ai/docs/)

### Muse Glimmer 30B
**Muse Glimmer 30B** (Meta; current; llm; released 2026-08; open weights) | ctx 131,072 | $0.3 in / $1.2 out per 1M tokens (USD) on OpenRouter; open weights free to self-host | OpenRouter: `meta/muse-glimmer-30b` | Hugging Face: https://huggingface.co/meta-models/Muse-Glimmer-30B — Released early Aug 2026 (exact day not verified). HF org is meta-models, not meta-llama. No first-party Meta API id verified.
  - Hugging Face — https://huggingface.co/meta-models/Muse-Glimmer-30B
  - OpenRouter: `meta/muse-glimmer-30b` — https://openrouter.ai/meta/muse-glimmer-30b
  - Pricing source: https://openrouter.ai/meta/muse-glimmer-30b
Notable capabilities:
  - Meta open weights under Apache 2.0: ~29.6B dense text+image model released Apache 2.0 (Llama used a custom community license), with llama.cpp / MLX / ExecuTorch integrations. (https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)
  - Local agents on one consumer GPU: Quantized to under 20GB for 24-32GB consumer GPUs/Macs; bundled DFlash drafter for speculative decoding gives ~3.1x speed-up on RTX 5090. (https://huggingface.co/meta-models/Muse-Glimmer-30B)
  - Agentic focus for its size: Optimized for multi-step reasoning, reliable tool use and failure recovery; Meta benchmarks it as competitive with Gemma4-31B and Qwen3.6-27B on agentic/coding evals. (https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)

### Omnilingual ASR
**Omnilingual ASR** (Meta; current; audio/speech; released 2025-11-10; open weights) | GitHub (fairseq2 checkpoints): `omniASR_LLM_7B_v2` | Hugging Face (demo space and dataset): https://huggingface.co/facebook — Open (Apache 2.0) suite: CTC and LLM-ASR models at 300M/1B/3B/7B, v2 checkpoints and 'Unlimited' long-audio LLM-ASR variants added December 2025, plus a 7B wav2vec 2.0 speech encoder and a corpus covering 350+ underserved languages. Checkpoints download via fairseq2 (e.g. https://dl.fbaipublicfiles.com/mms/omniASR-LLM-7B-v2.pt). Successor to MMS. The 'first' claim is Meta's ('never previously supported by any ASR model').
  - GitHub (fairseq2 checkpoints): `omniASR_LLM_7B_v2` — https://github.com/facebookresearch/omnilingual-asr
  - Hugging Face (demo space and dataset) — https://huggingface.co/facebook
Notable capabilities:
  - [FIRST] ASR for 1,600+ languages: Transcribes 1,600+ languages, ~500 of them never before supported by any ASR system (Whisper covers 99). (https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/)
  - Zero-shot in-context language extension: omniASR_LLM_7B_ZS transcribes new languages from a few paired audio-text examples at inference, extending potential coverage to 5,400+ languages. (https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/)

### Llama 4 Maverick (17B-128E)
**Llama 4 Maverick (17B-128E)** (Meta; legacy; multimodal; released 2025-04-05; open weights) | ctx 1,000,000 | AWS Bedrock: `meta.llama4-maverick-17b-instruct-v1:0`; OpenRouter: `meta-llama/llama-4-maverick` | Hugging Face: https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct; Web app: https://meta.ai — FP8 repo meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8. Bedrock max output 8K. No first-party pay-as-you-go pricing verified. Superseded at Meta by closed Muse Spark and open Muse Glimmer.
  - Hugging Face — https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct
  - AWS Bedrock: `meta.llama4-maverick-17b-instruct-v1:0` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-maverick-17b-instruct.html)
  - OpenRouter: `meta-llama/llama-4-maverick` — https://openrouter.ai/meta-llama/llama-4-maverick
  - Web app — https://meta.ai
Notable capabilities:
  - [FIRST] First natively multimodal Llama (early fusion): Llama 4 were the first Llama models with native multimodality via early fusion of text and vision tokens. (https://ai.meta.com/blog/llama-4-multimodal-intelligence/)
  - 400B-total MoE on one H100 host: 17B active / 128 experts / ~400B total; runs on a single H100 host. (https://ai.meta.com/blog/llama-4-multimodal-intelligence/)
  - LMArena experimental-variant controversy (found after launch): Launch LMArena Elo 1417 came from an experimental chat-tuned variant, not the released weights, drawing criticism. (https://en.wikipedia.org/wiki/Llama_(language_model))

### Llama 4 Scout (17B-16E)
**Llama 4 Scout (17B-16E)** (Meta; legacy; multimodal; released 2025-04-05; open weights) | AWS Bedrock: `meta.llama4-scout-17b-instruct-v1:0`; OpenRouter: `meta-llama/llama-4-scout` | Hugging Face: https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct; Web app: https://meta.ai — Context: 10M per Meta; provider limits vary (not listed as context_window). Knowledge cutoff Aug 2024 per Meta model card (not re-verified today). No first-party pricing verified.
  - Hugging Face — https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct
  - AWS Bedrock: `meta.llama4-scout-17b-instruct-v1:0` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-scout-17b-instruct.html)
  - OpenRouter: `meta-llama/llama-4-scout` — https://openrouter.ai/meta-llama/llama-4-scout
  - Web app — https://meta.ai
Notable capabilities:
  - [FIRST] 10M-token context (claimed): Meta advertised an 'industry-leading' 10M-token context via the iRoPE architecture; hosted providers typically serve far less (e.g. ~1.3M on OpenRouter). (https://ai.meta.com/blog/llama-4-multimodal-intelligence/)
  - Single-H100 multimodal MoE: 17B active / 16 experts / 109B total; fits one H100 with Int4 quantization. (https://ai.meta.com/blog/llama-4-multimodal-intelligence/)


## Microsoft

### Phi-4-Reasoning-Vision-15B
**Phi-4-Reasoning-Vision-15B** (Microsoft; current; multimodal; released 2026-03-04; open weights) | ctx 16,384 | Hugging Face: https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B; Microsoft Foundry: https://aka.ms/Phi-4-r-v-foundry — Newest Phi model found (Mar 2026). Foundry model id and pricing not verified. Microsoft MAI models (MAI-Image-2/2.5, MAI-Voice-2, MAI-Transcribe-2, MAI-Thinking-1) are in Foundry but not covered by a file here.
  - Hugging Face — https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B
  - Microsoft Foundry — https://aka.ms/Phi-4-r-v-foundry (docs: https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-phi-4-reasoning-vision-to-microsoft-foundry/4499154)
Notable capabilities:
  - Hybrid think / no-think vision reasoning: Automatically chooses direct answers for perception tasks and long chain-of-thought only for math/science/diagram problems. (https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B)
  - GUI grounding for computer-use agents: Dynamic-resolution SigLIP-2 encoder (up to 3,600 visual tokens) with strengths in GUI grounding for computer-use agents. (https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B)

### VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS)
**VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS)** (Microsoft; current; audio/speech; released 2025-08-25; open weights) | Hugging Face: `microsoft/VibeVoice-ASR`; Hugging Face (Transformers format): `microsoft/VibeVoice-ASR-HF`; Hugging Face: `microsoft/VibeVoice-ASR-BitNet`; Hugging Face: `microsoft/VibeVoice-Realtime-0.5B`; Hugging Face (streaming ASR, 7B repo; 9B params incl. decoder): `microsoft/VibeVoice-ASR-Streaming-7B`; Hugging Face (streaming ASR, small): `microsoft/VibeVoice-ASR-Streaming-1.5B` | GitHub: https://github.com/microsoft/VibeVoice — Open-source voice research family from Microsoft (MIT). Timeline: TTS 2025-08-25 (code pulled 2025-09-05), Realtime-0.5B streaming TTS (~300 ms first audio) 2025-12-03, ASR 2026-01-21, Transformers integration 2026-03, Foundry Labs 2026-03-12, ASR-BitNet 2026-07-23, ASR-Streaming (10 languages, hotwords, speaker attribution) announced 2026-09-03; HF repos microsoft/VibeVoice-ASR-Streaming-7B and -1.5B created 2026-09-02 (verified 2026-09-29). Monthly downloads to 2026-09-29: VibeVoice-ASR ~734k, VibeVoice-1.5B ~717k. Separate from Microsoft's proprietary MAI-Voice/MAI-Transcribe.
  - Hugging Face: `microsoft/VibeVoice-ASR` — https://huggingface.co/microsoft/VibeVoice-ASR
  - Hugging Face (Transformers format): `microsoft/VibeVoice-ASR-HF` — https://huggingface.co/microsoft/VibeVoice-ASR-HF
  - Hugging Face: `microsoft/VibeVoice-ASR-BitNet` — https://huggingface.co/microsoft/VibeVoice-ASR-BitNet
  - Hugging Face: `microsoft/VibeVoice-Realtime-0.5B` — https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B
  - Hugging Face (streaming ASR, 7B repo; 9B params incl. decoder): `microsoft/VibeVoice-ASR-Streaming-7B` — https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B
  - Hugging Face (streaming ASR, small): `microsoft/VibeVoice-ASR-Streaming-1.5B` — https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-1.5B
  - GitHub — https://github.com/microsoft/VibeVoice
Notable capabilities:
  - 60-minute single-pass ASR with diarization: VibeVoice-ASR (~9B params incl. Qwen2-based decoder) transcribes up to 60 min in one pass with who/when/what structured output, hotwords and 50+ languages with code-switching. (https://huggingface.co/microsoft/VibeVoice-ASR)
  - CPU-only realtime ASR (found after launch): VibeVoice-ASR-BitNet (2026-07-23) compresses the model 4.62 GB -> 1.58 GB and runs faster than real time on 3 CPU threads (1.6-2.3x faster than Whisper.cpp). (https://huggingface.co/microsoft/VibeVoice-ASR-BitNet)
  - Streaming speaker-attributed ASR (found after launch): VibeVoice-ASR-Streaming (7B and 1.5B repos, uploaded 2026-09-02) transcribes live audio with speaker attribution (who said what) and custom hotwords in 10 languages (zh, en, fr, de, it, ja, ko, pt, ru, es); MIT license. (https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B)
  - Long-form multi-speaker TTS (withdrawn) (found after launch): Original VibeVoice-TTS (1.5B/7B) generated up to 90 min with 4 speakers; Microsoft removed the TTS code on 2025-09-05 over responsible-AI misuse concerns. (https://github.com/microsoft/VibeVoice)

### MAI-Transcribe-2
**MAI-Transcribe-2** (Microsoft; preview; audio/speech; released 2026-09-03) | Azure Speech in Microsoft Foundry (Fast Transcription API, enhancedMode): `MAI-Transcribe-2`; Azure Speech (previous version): `MAI-Transcribe-1.5`; Azure Voice Live (input transcription): `MAI-Transcribe-2`; OpenRouter: `microsoft/mai-transcribe-2` | Web app (MAI Playground): https://playground.microsoft.ai/ — Public preview in Azure Speech. MAI-Transcribe-1.5 (Build 2026-06-02, 43 languages, $0.36/hr) remains available; MAI-Transcribe-1 deprecated 2026-08-20. Standard (post-promo) price not published. Input WAV/MP3/FLAC. Model card: https://microsoft.ai/pdf/MAI-Transcribe-2-Model-Card.pdf. Benchmarks are Microsoft-reported.
  - Azure Speech in Microsoft Foundry (Fast Transcription API, enhancedMode): `MAI-Transcribe-2` — https://{resource}.cognitiveservices.azure.com/speechtotext/transcriptions:transcribe?api-version=2025-10-15 (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe)
  - Azure Speech (previous version): `MAI-Transcribe-1.5` (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe)
  - Azure Voice Live (input transcription): `MAI-Transcribe-2` (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe)
  - OpenRouter: `microsoft/mai-transcribe-2` — https://openrouter.ai/api/v1/audio/transcriptions
  - Web app (MAI Playground) — https://playground.microsoft.ai/
  - Pricing source: https://microsoft.ai/models/mai-transcribe-2/
Notable capabilities:
  - #1 on FLEURS across 60 languages (claimed): Microsoft reports 5.2% average WER over 60 FLEURS languages (3.4% on top-25) and #2 on the Artificial Analysis WER leaderboard. (https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/)
  - Very fast batch transcription: Claims ~10x faster than GPT-Transcribe (1 hour of audio in ~10 s), 7x vs Scribe v2, 5x vs Gemini 3.5. (https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/)
  - Diarization, word timestamps, keyword biasing, clean/verbatim styles: New in v2: speaker diarization, word-level timestamps, phrase-list biasing, code-switching (e.g. Hinglish) and verbatim vs clean transcripts. (https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe)

### MAI-Voice-2 / MAI-Voice-2-Flash
**MAI-Voice-2 / MAI-Voice-2-Flash** (Microsoft; preview; audio/speech; released 2026-06-02) | Azure Speech in Microsoft Foundry (SSML voice name): `en-US-Harper:MAI-Voice-2`; Azure Speech in Microsoft Foundry (SSML voice name): `en-US-Harper:MAI-Voice-2-Flash`; Azure Voice Live (TTS output): `MAI-Voice-2-Flash`; OpenRouter: `microsoft/mai-voice-2`; OpenRouter: `microsoft/mai-voice-2-flash` | Web app (MAI Playground): https://playground.microsoft.ai/ — Launched at Build 2026-06-02 (MAI-Voice-2); Flash followed 2026-07-23 (date per secondary sources). Both public preview in Azure Speech. Languages include en-US/AU, de, fr, es-ES/MX, pt-BR/PT, it, ko, zh-CN, tr, ru, th, nl, ro, hu, hi. Also used in Copilot (Audio Expressions). Predecessor MAI-Voice-1 no longer listed on the MAI-Voice docs page. Also on Fireworks and Baseten (ids not verified).
  - Azure Speech in Microsoft Foundry (SSML voice name): `en-US-Harper:MAI-Voice-2` — https://{region}.tts.speech.microsoft.com/cognitiveservices/v1 (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices)
  - Azure Speech in Microsoft Foundry (SSML voice name): `en-US-Harper:MAI-Voice-2-Flash` (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices)
  - Azure Voice Live (TTS output): `MAI-Voice-2-Flash` (docs: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-how-to)
  - OpenRouter: `microsoft/mai-voice-2` — https://openrouter.ai/api/v1/audio/speech
  - OpenRouter: `microsoft/mai-voice-2-flash`
  - Web app (MAI Playground) — https://playground.microsoft.ai/
  - Pricing source: https://microsoft.ai/models/mai-voice-2/
Notable capabilities:
  - Gated instant voice cloning: Matches a consented reference voice from a 5-60 s clip without training; only approved (Limited Access) licensed voices can be synthesized. (https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices)
  - SSML emotion/style control: mstts:express-as styles (angry, fearful, joyful, whispering, shouting, etc.) with styledegree, across 15 languages / 18 locales. (https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices)
  - Low-latency Flash tier (found after launch): MAI-Voice-2-Flash (public preview from 2026-07-23) targets voice agents/IVR; Microsoft quotes ~225 ms latency vs ~1 s for MAI-Voice-2 (for a 45 s clip). (https://microsoft.ai/models/mai-voice-2/)

### Phi-4 (14B)
**Phi-4 (14B)** (Microsoft; legacy; llm; released 2024-12-12; open weights) | ctx 16,384 | $0.07 in / $0.14 out per 1M tokens (USD) on OpenRouter; self-hosting free | Azure AI Foundry: `Phi-4`; OpenRouter: `microsoft/phi-4` | Hugging Face: https://huggingface.co/microsoft/phi-4 — No Phi-5 found on Hugging Face as of 2026-09-29 (microsoft org). Foundry model name not re-verified today. Siblings: microsoft/Phi-4-mini-instruct, microsoft/Phi-4-reasoning-plus, microsoft/Phi-4-multimodal-instruct.
  - Hugging Face — https://huggingface.co/microsoft/phi-4
  - Azure AI Foundry: `Phi-4` — https://azure.microsoft.com/en-us/products/phi
  - OpenRouter: `microsoft/phi-4` — https://openrouter.ai/microsoft/phi-4
  - Pricing source: https://openrouter.ai/microsoft/phi-4
Notable capabilities:
  - Synthetic-data small model: 14B dense model trained on 9.8T tokens heavy in curated synthetic data, prioritizing reasoning over scale (84.8 MMLU, 80.4 MATH). (https://huggingface.co/microsoft/phi-4)
  - Reasoning derivatives (found after launch): Base for Phi-4-reasoning, Phi-4-reasoning-plus, Phi-4-mini(-reasoning/-flash-reasoning) and Phi-4-multimodal-instruct open models. (https://huggingface.co/microsoft)


## Midjourney

### Midjourney V8.2
**Midjourney V8.2** (Midjourney; current; image-gen; released 2026-07-24) | Web app: https://www.midjourney.com; Discord: https://discord.gg/midjourney — No official public API (web app/Discord only; subscription). V8 alpha 2026-03-17, V8.1 2026-04-14 (default from 2026-06-10), V8.2 2026-07-24 - reportedly now default (not confirmed on an official page). Select with --v 8.2 (syntax per docs; docs page blocked). Pricing not verified.
  - Web app — https://www.midjourney.com
  - Discord — https://discord.gg/midjourney
Notable capabilities:
  - Instruction-based edit model (found after launch): V8.2 edit model (Aug 2026) edits images from plain instructions, takes up to 4 image references (replacing Omni Reference / Character Reference / Retexture) and does inpainting/outpainting. (https://updates.midjourney.com/edit-model-for-v8/)
  - Improved personalization: V8.2 release focused on aesthetics and personalization profiles that better learn a user's taste from image ratings. (https://updates.midjourney.com/version-8-2/)
  - Rewritten V8 core with native 2K and better text: V8 line (alpha 2026-03-17, V8.1 2026-04-14) was rebuilt from scratch: much faster jobs, HD/2K output, better prompt following and in-image text. (https://updates.midjourney.com/v8-alpha/)


## MiniMax

### MiniMax H3
**MiniMax H3** (MiniMax; current; video-gen; released 2026-07-31; open weights) | MiniMax API (Video Generation V2): `MiniMax-H3`; MiniMax API (fast variant): `MiniMax-H3-Max` | Hugging Face: https://huggingface.co/MiniMaxAI/MiniMax-H3 — Replaces Hailuo 2.3 / 2.3-Fast / 02 (now legacy: e.g. MiniMax-Hailuo-2.3 0.28 USD per 768P 6s clip). Modes: T2V, I2V, first/last frame, multimodal reference; 4-15 s, 24 fps. Open release is full-attention only.
  - MiniMax API (Video Generation V2): `MiniMax-H3` — https://api.minimax.io (docs: https://platform.minimax.io/docs/api-reference/video-generation-v2-create)
  - MiniMax API (fast variant): `MiniMax-H3-Max` — https://api.minimax.io (docs: https://platform.minimax.io/docs/api-reference/video-generation-v2-create)
  - Hugging Face — https://huggingface.co/MiniMaxAI/MiniMax-H3
  - Pricing source: https://platform.minimax.io/docs/guides/pricing-paygo
Notable capabilities:
  - Open omni-modal video model with native audio: Understands mixed text/image/video/audio context and generates video with native stereo audio, up to 2K and 15 s. (https://huggingface.co/MiniMaxAI/MiniMax-H3)
  - H3-Context-IR prompt pipeline: Hosted system turns free-form multimodal instructions into a structured intermediate representation before generation (API-only, not open-sourced). (https://huggingface.co/MiniMaxAI/MiniMax-H3)
  - 768P to 2K regeneration: H3-Regenerate-2K re-renders a 768P result with the original context into 2K (0.05 USD/s). (https://platform.minimax.io/docs/guides/pricing-paygo)

### MiniMax Music 3.0
**MiniMax Music 3.0** (MiniMax; current; music; released 2026-07-16; open weights) | MiniMax API (existing paying users only since 2026-08-20): `music-3.0` | Hugging Face: https://huggingface.co/MiniMaxAI/MiniMax-Music3; GitHub: https://github.com/MiniMax-AI/MiniMax-Music3; Web app (MiniMax Audio): https://www.minimax.io/audio — music-3.0 shipped on the MiniMax API on 2026-07-16 (release notes); open weights published 2026-08-13. Earlier API models: music-2.6 (Apr 2026, covers), music-cover, music-2.5 (Jan 2026), music-2.0 (legacy). On 2026-08-20 MiniMax stopped offering the paid Music and Lyrics Generation APIs to new users and points them to MiniMax Audio or the open model. Demonstrated with English and Mandarin lyrics; no third-party benchmark vs Suno found.
  - MiniMax API (existing paying users only since 2026-08-20): `music-3.0` (docs: https://platform.minimax.io/docs/api-reference/music-generation)
  - Hugging Face — https://huggingface.co/MiniMaxAI/MiniMax-Music3
  - GitHub — https://github.com/MiniMax-AI/MiniMax-Music3
  - Web app (MiniMax Audio) — https://www.minimax.io/audio
  - Pricing source: https://platform.minimax.io/docs/guides/pricing-paygo
Notable capabilities:
  - Open-weights full songs up to ~5 minutes in one pass: Composes, arranges, performs and produces a complete song (vocals + arrangement) up to about five minutes from lyrics with section tags and a structured caption; 32 kHz 16-bit stereo WAV. (https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)
  - Hierarchical Global/Local LLM with continuous hidden-state synthesis: 8B Global LLM (initialized from Qwen3.5-8B) for long-range structure + 0.6B Local LLM for frame-level acoustics, rendered by a 2.4B flow-matching module and 123M Flow-VAE instead of discrete token decoding. (https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model)
  - Consumer-GPU inference: 24 GB+ VRAM recommended; runs on 8 GB with CPU offloading; diffusers modular pipeline and ComfyUI support (Comfy-Org/MiniMax-Music-3). (https://huggingface.co/MiniMaxAI/MiniMax-Music3)

### MiniMax-M3
**MiniMax-M3** (MiniMax; current; reasoning-llm; released 2026-06-01; open weights) | ctx 1,000,000 | $0.3 in / $1.2 out per 1M tokens (USD), standard tier, input <=512K (after permanent 50% discount); >512K input: 0.60/2.40/0.12. Priority tier 1.5x | MiniMax API (Anthropic format): `MiniMax-M3`; MiniMax API (OpenAI format): `MiniMax-M3`; OpenRouter: `minimax/minimax-m3` | Hugging Face: https://huggingface.co/MiniMaxAI/MiniMax-M3 — MiniMax flagship LLM. OpenAI-format responses include <think> content that must be preserved across turns. MiniMax-M3.1-Flash-Preview (1M, tunable thinking) exists but only via Token Plan/MiniMax Code. Max output and knowledge cutoff not verified.
  - MiniMax API (Anthropic format): `MiniMax-M3` — https://api.minimax.io/anthropic (docs: https://platform.minimax.io/docs/api-reference/text-anthropic-api)
  - MiniMax API (OpenAI format): `MiniMax-M3` — https://api.minimax.io/v1 (docs: https://platform.minimax.io/docs/api-reference/text-openai-api)
  - OpenRouter: `minimax/minimax-m3` — https://openrouter.ai/minimax/minimax-m3
  - Hugging Face — https://huggingface.co/MiniMaxAI/MiniMax-M3
  - Pricing source: https://platform.minimax.io/docs/guides/pricing-paygo
Notable capabilities:
  - MiniMax Sparse Attention (MSA): New sparse attention for million-token contexts: 9x prefill and 15x decode speed-up vs M2 at 1M context, ~1/20 per-token compute. (https://huggingface.co/MiniMaxAI/MiniMax-M3)
  - Native multimodality from step one: Mixed text/image/video training from the start of pre-training (~428B total / ~23B active). (https://huggingface.co/MiniMaxAI/MiniMax-M3)
  - Three reasoning modes: thinking parameter selects among three reasoning modes; interleaved thinking with tool use. (https://huggingface.co/MiniMaxAI/MiniMax-M3)

### MiniMax-M2.7
**MiniMax-M2.7** (MiniMax; current; reasoning-llm; released 2026-03-18; open weights) | ctx 204,800 | $0.3 in / $1.2 out per 1M tokens (USD); MiniMax-M2.7-highspeed: 0.6 / 2.4 | MiniMax API (Anthropic format): `MiniMax-M2.7`; MiniMax API (OpenAI format): `MiniMax-M2.7`; MiniMax API (fast): `MiniMax-M2.7-highspeed`; OpenRouter: `minimax/minimax-m2.7` | Hugging Face: https://huggingface.co/MiniMaxAI/MiniMax-M2.7 — Text-only predecessor of M3, still a current API model; highspeed variant ~100 tok/s vs ~60. M2.5/M2.1/M2 are legacy (same $0.3/$1.2 price).
  - MiniMax API (Anthropic format): `MiniMax-M2.7` — https://api.minimax.io/anthropic (docs: https://platform.minimax.io/docs/api-reference/text-anthropic-api)
  - MiniMax API (OpenAI format): `MiniMax-M2.7` — https://api.minimax.io/v1 (docs: https://platform.minimax.io/docs/api-reference/text-openai-api)
  - MiniMax API (fast): `MiniMax-M2.7-highspeed` — https://api.minimax.io/v1 (docs: https://platform.minimax.io/docs/api-reference/text-openai-api)
  - OpenRouter: `minimax/minimax-m2.7` — https://openrouter.ai/minimax/minimax-m2.7
  - Hugging Face — https://huggingface.co/MiniMaxAI/MiniMax-M2.7
  - Pricing source: https://platform.minimax.io/docs/guides/pricing-paygo
Notable capabilities:
  - Participates in its own evolution: MiniMax calls it its first model deeply participating in its own development ('recursive self-improvement'). (https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
  - Agent harness building: Builds complex agent harnesses using Agent Teams, Skills and dynamic tool search; aimed at professional office delivery. (https://huggingface.co/MiniMaxAI/MiniMax-M2.7)

### MiniMax Speech 2.8 (HD / Turbo)
**MiniMax Speech 2.8 (HD / Turbo)** (MiniMax; current; audio/speech; released 2026-01-23) | MiniMax API (T2A HTTP / WebSocket / async): `speech-2.8-hd`; MiniMax API: `speech-2.8-turbo` | Web app (MiniMax Audio): https://www.minimax.io/audio — speech-2.6 and speech-02 are legacy at the same prices. MiniMax also offers ASR (0.38 USD/hour). Music: music-3.0 API closed to new users from 2026-08-20; open weights MiniMax-Music3 on HF.
  - MiniMax API (T2A HTTP / WebSocket / async): `speech-2.8-hd` — https://api.minimax.io (docs: https://platform.minimax.io/docs/api-reference/speech-t2a-http)
  - MiniMax API: `speech-2.8-turbo` — https://api.minimax.io (docs: https://platform.minimax.io/docs/api-reference/speech-t2a-http)
  - Web app (MiniMax Audio) — https://www.minimax.io/audio
  - Pricing source: https://platform.minimax.io/docs/guides/pricing-paygo
Notable capabilities:
  - Sound tags: Natural sound tags (non-verbal cues) in ultra-realistic HD speech. (https://platform.minimax.io/docs/release-notes/models)
  - 40 languages, 7 emotions: 40 languages plus specified dialects, 7 emotions; rapid voice cloning and text-described voice design. (https://platform.minimax.io/docs/guides/models-intro)
  - Streaming and long-form modes: Sync HTTP, WebSocket and bidirectional streaming (pipe LLM tokens straight to speech), plus async jobs up to 1M characters. (https://platform.minimax.io/docs/guides/pricing-paygo)


## Mistral AI

### Mistral OCR 4.1
**Mistral OCR 4.1** (Mistral AI; current; multimodal; released 2026-07-16) | Mistral API: `mistral-ocr-4-1` — Aliases mistral-ocr-4 and mistral-ocr-latest point to 4.1. Powers Mistral Document AI.
  - Mistral API: `mistral-ocr-4-1` — https://api.mistral.ai/v1/ocr (docs: https://docs.mistral.ai/models/model-cards/ocr-4-1)
  - Pricing source: https://docs.mistral.ai/models/model-cards/ocr-4-1
Notable capabilities:
  - Paragraph-level bounding boxes with confidence: Native paragraph-level bbox extraction, structural block labels and block-level confidence scores. (https://docs.mistral.ai/models/model-cards/ocr-4-1)
  - Structured annotations: Schema-driven document annotation priced separately ($5 / 1,000 annotated pages); batch via /v1/batch. (https://docs.mistral.ai/models/model-cards/ocr-4-1)

### Mistral Medium 3.5
**Mistral Medium 3.5** (Mistral AI; current; multimodal; released 2026-04-28; open weights) | ctx 256,000 | $1.5 in / $7.5 out per 1M tokens (USD) | Mistral API: `mistral-medium-3-5`; OpenRouter: `mistralai/mistral-medium-3-5` | Hugging Face: https://huggingface.co/mistralai/Mistral-Medium-3.5-128B; Web app: https://chat.mistral.ai — Alias mistral-medium-latest (version v26.04). Official card lists 2 more aliases not verified. Batch API supported (OpenRouter batch $0.75/$3.75). Knowledge cutoff not published.
  - Mistral API: `mistral-medium-3-5` — https://api.mistral.ai/v1/chat/completions (docs: https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04)
  - OpenRouter: `mistralai/mistral-medium-3-5` — https://openrouter.ai/mistralai/mistral-medium-3-5
  - Hugging Face — https://huggingface.co/mistralai/Mistral-Medium-3.5-128B
  - Web app — https://chat.mistral.ai
  - Pricing source: https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04
Notable capabilities:
  - One model replacing Devstral 2 and Magistral: Frontier-class multimodal model for agentic and coding use; Mistral names it the replacement for deprecated Devstral 2 (deprecated 2026-05-22). (https://docs.mistral.ai/models/model-cards/devstral-2-25-12)
  - Open-weight 128B dense with vision: 128B dense weights on Hugging Face under a modified MIT license, 256K context, built-in tools and Agents/Conversations API support. (https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04)

### Voxtral TTS
**Voxtral TTS** (Mistral AI; current; audio/speech; released 2026-03-23; open weights) | Mistral API: `voxtral-tts-2603`; Hugging Face: `mistralai/Voxtral-4B-TTS-2603` | Web app: https://chat.mistral.ai — Mistral's first TTS model. 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, Arabic. Weights are CC BY-NC 4.0 (non-commercial); commercial use via API. Docs model-card page shows id voxtral-tts-2603 on the overview (a 'voxtral-mini-tts-2603' alias also appears on the card).
  - Mistral API: `voxtral-tts-2603` — https://api.mistral.ai/v1/audio/speech (docs: https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03)
  - Hugging Face: `mistralai/Voxtral-4B-TTS-2603` — https://huggingface.co/mistralai/Voxtral-4B-TTS-2603
  - Web app — https://chat.mistral.ai
  - Pricing source: https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03
Notable capabilities:
  - Zero-shot voice cloning from ~3 s: Clones a voice (accent, fillers, rhythm) from a few seconds of reference audio without a transcript; 68.4% human-preference win rate vs ElevenLabs Flash v2.5 on multilingual cloning (Mistral-reported). (https://mistral.ai/news/voxtral-tts)
  - Open-weight 4B TTS with low latency: 3.4B decoder + 390M flow-matching acoustic transformer + 300M codec; ~70 ms model latency (~90 ms time-to-first-audio via API), RTF ~9.7x, up to 2 min native generation. (https://mistral.ai/news/voxtral-tts)

### Mistral Small 4
**Mistral Small 4** (Mistral AI; current; reasoning-llm; released 2026-03-16; open weights) | ctx 256,000 | $0.15 in / $0.6 out per 1M tokens (USD) | Mistral API: `mistral-small-2603`; OpenRouter: `mistralai/mistral-small-2603` | Hugging Face: https://huggingface.co/mistralai/Mistral-Small-4-119B-2603; Web app: https://chat.mistral.ai — Alias mistral-small-latest (v26.03). Announced Mar 16, 2026.
  - Mistral API: `mistral-small-2603` — https://api.mistral.ai/v1/chat/completions (docs: https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03)
  - OpenRouter: `mistralai/mistral-small-2603` — https://openrouter.ai/mistralai/mistral-small-2603
  - Hugging Face — https://huggingface.co/mistralai/Mistral-Small-4-119B-2603
  - Web app — https://chat.mistral.ai
  - Pricing source: https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03
Notable capabilities:
  - Instruct + reasoning + coding unified: First Mistral model unifying Magistral (reasoning), Pixtral (multimodal) and Devstral (agentic coding) in one model; reasoning_effort none/high per request. (https://mistral.ai/news/mistral-small-4/)
  - 119B MoE with ~6.5B active: 119B total / 6.5B active parameters, vision input, 256K context at $0.15/$0.6. (https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03)

### Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime)
**Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime)** (Mistral AI; current; audio/speech; released 2026-02-04; open weights) | Mistral API (batch): `voxtral-mini-2602`; Mistral API (realtime): `voxtral-mini-transcribe-realtime-2602`; Hugging Face (Realtime, open weights): `mistralai/Voxtral-Mini-4B-Realtime-2602` | Web app: https://chat.mistral.ai — 13 languages (en, zh, hi, es, ar, fr, pt, ru, de, ja, ko, it, nl). Batch model is API-only ('Premier' license); Realtime has open weights. Replaced voxtral-mini-2507 / Voxtral Mini Transcribe (deprecated 2026-02-27, retired 2026-05-31). Tech report arXiv 2602.11298. Accuracy claims are Mistral's.
  - Mistral API (batch): `voxtral-mini-2602` — https://api.mistral.ai/v1/audio/transcriptions (docs: https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-26-02)
  - Mistral API (realtime): `voxtral-mini-transcribe-realtime-2602` (docs: https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-realtime-26-02)
  - Hugging Face (Realtime, open weights): `mistralai/Voxtral-Mini-4B-Realtime-2602` — https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602
  - Web app — https://chat.mistral.ai
  - Pricing source: https://mistral.ai/news/voxtral-transcribe-2
Notable capabilities:
  - Open-weight realtime ASR under 200 ms: Voxtral Realtime (4B, Apache 2.0) reaches sub-200 ms latency; at 480 ms delay Mistral reports 1-2% WER. (https://mistral.ai/news/voxtral-transcribe-2)
  - Cheap batch transcription with diarization: Mini Transcribe V2: ~4% WER on FLEURS at $0.003/min with speaker diarization, word timestamps, context biasing (up to 100 terms) and audio up to 3 hours. (https://mistral.ai/news/voxtral-transcribe-2)

### Mistral Large 3
**Mistral Large 3** (Mistral AI; current; multimodal; released 2025-12-02; open weights) | ctx 256,000 | $0.5 in / $1.5 out per 1M tokens (USD) | Mistral API: `mistral-large-2512`; AWS Bedrock: `mistral.mistral-large-3-675b-instruct`; OpenRouter: `mistralai/mistral-large-2512` | Hugging Face: https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512; Web app: https://chat.mistral.ai — Alias mistral-large-latest (v25.12). Still GA; for coding/agents Mistral now points to Medium 3.5.
  - Mistral API: `mistral-large-2512` — https://api.mistral.ai/v1/chat/completions (docs: https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12)
  - AWS Bedrock: `mistral.mistral-large-3-675b-instruct` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-mistral-ai-mistral-large-3.html)
  - OpenRouter: `mistralai/mistral-large-2512` — https://openrouter.ai/mistralai/mistral-large-2512
  - Hugging Face — https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512
  - Web app — https://chat.mistral.ai
  - Pricing source: https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12
Notable capabilities:
  - 675B open-weight MoE under Apache 2.0: Granular mixture-of-experts with 41B active / 675B total parameters, fully Apache 2.0. (https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12)
  - Very low price for size: $0.5 / $1.5 per 1M tokens with 256K context and vision - cheaper than Mistral Medium 3.5. (https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12)

### Codestral 25.08
**Codestral 25.08** (Mistral AI; current; code; released 2025-07-30) | ctx 128,000 | $0.3 in / $0.9 out per 1M tokens (USD) | Mistral API (FIM): `codestral-2508`; Mistral API (chat): `codestral-latest`; OpenRouter: `mistralai/codestral-2508` — Alias codestral-latest. Mistral's current code-completion model (Premier). OpenRouter lists 256K context; Mistral card says 128K.
  - Mistral API (FIM): `codestral-2508` — https://api.mistral.ai/v1/fim/completions (docs: https://docs.mistral.ai/models/model-cards/codestral-25-08)
  - Mistral API (chat): `codestral-latest` — https://api.mistral.ai/v1/chat/completions
  - OpenRouter: `mistralai/codestral-2508` — https://openrouter.ai/mistralai/codestral-2508
  - Pricing source: https://docs.mistral.ai/models/model-cards/codestral-25-08
Notable capabilities:
  - Low-latency fill-in-the-middle: Specialized for high-frequency FIM/autocomplete with a dedicated FIM endpoint, plus predicted outputs. (https://docs.mistral.ai/models/model-cards/codestral-25-08)
  - Predicted outputs and prefix mode: Supports predicted outputs (fast edits of known code) and assistant prefix, plus function calling and structured outputs. (https://docs.mistral.ai/models/model-cards/codestral-25-08)

### Voxtral Small
**Voxtral Small** (Mistral AI; current; multimodal; released 2025-07; open weights) | Mistral API: `voxtral-small-2507`; Hugging Face: `mistralai/Voxtral-Small-24B-2507` — Still listed as active (v25.07) on Mistral's models overview on 2026-09-29; its small siblings voxtral-mini-2507 and Voxtral Mini Transcribe 25.07 were retired 2026-05-31. Pricing and exact release day not re-verified (July 2025 launch).
  - Mistral API: `voxtral-small-2507` — https://api.mistral.ai/v1/chat/completions (docs: https://docs.mistral.ai/models/overview)
  - Hugging Face: `mistralai/Voxtral-Small-24B-2507` — https://huggingface.co/mistralai/Voxtral-Small-24B-2507
Notable capabilities:
  - Audio-understanding chat model: Mistral's first model with audio input for instruct use (Q&A, summarization, function calling from voice) on top of transcription. (https://docs.mistral.ai/models/overview)

### Robostral Navigate
**Robostral Navigate** (Mistral AI; preview; robotics; released 2026-07-08) | Mistral AI (contact sales / partners; no public API id or weights found): https://mistral.ai/news/robostral-navigate/ — Mistral's first robotics model; built in-house without an existing open VLM. Outputs navigation actions. Access appears to be via Mistral's team ('talk with our team'); status set to preview.
  - Mistral AI (contact sales / partners; no public API id or weights found) — https://mistral.ai/news/robostral-navigate/ (docs: https://arxiv.org/abs/2607.20785)
Notable capabilities:
  - Single-RGB-camera vision-language navigation: 8B model navigates buildings from one RGB camera plus language instructions (no LiDAR/depth); R2R-CE success 79.4% val-seen, 76.6% val-unseen (+9.7 pts over best single-camera method, +4.5 over depth/multi-camera systems). (https://mistral.ai/news/robostral-navigate/)
  - Sim-only training, embodiment-agnostic: Trained in simulation (~2.4M trajectories across 350k scenes per Mistral's page), with prefix caching (22x fewer training tokens) and online RL (CISPO, +3.2 pts); works on wheeled, legged and flying robots. (https://mistral.ai/news/robostral-navigate/)


## Moonshot AI

### Kimi K3
**Kimi K3** (Moonshot AI; current; reasoning-llm; released 2026-07-16; open weights) | ctx 1,048,576 | $3 in / $15 out per 1M tokens (USD) | Kimi API (Moonshot): `kimi-k3`; Alibaba Cloud Model Studio: `kimi-k3`; OpenRouter: `moonshotai/kimi-k3` | Hugging Face: https://huggingface.co/moonshotai/Kimi-K3; Web app: https://www.kimi.com — Moonshot flagship. API unlocked after a minimum $1 top-up. Chat Completions, Responses and Anthropic-compatible Messages supported. Max output and knowledge cutoff not verified. Docs moved to platform.kimi.ai (platform.moonshot.ai still serves).
  - Kimi API (Moonshot): `kimi-k3` — https://api.moonshot.ai/v1 (docs: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
  - Alibaba Cloud Model Studio: `kimi-k3` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - OpenRouter: `moonshotai/kimi-k3` — https://openrouter.ai/moonshotai/kimi-k3
  - Hugging Face — https://huggingface.co/moonshotai/Kimi-K3
  - Web app — https://www.kimi.com
  - Pricing source: https://platform.kimi.ai/docs/pricing/chat
Notable capabilities:
  - [FIRST] First open 3T-class model: 2.8T-parameter MoE (16 of 896 experts active) - Moonshot's claim: the first open model at this scale; weights released after launch (promised by 2026-07-27). (https://www.kimi.com/blog/kimi-k3)
  - Kimi Delta Attention + Attention Residuals: Hybrid linear attention (KDA) and AttnRes; ~2.5x the scaling efficiency of K2 per Moonshot. (https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
  - Native vision with 1M context: Native visual understanding (image and video) and a 1,048,576-token window; strong at coding tasks that use screenshots/visual feedback (games, frontend, CAD). (https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
  - Always-on thinking with effort control: Thinking cannot be disabled; reasoning_effort low/high/max (default max). (https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)

### Kimi K2.7 Code
**Kimi K2.7 Code** (Moonshot AI; current; code; released 2026-06; open weights) | ctx 262,144 | $0.95 in / $4 out per 1M tokens (USD); kimi-k2.7-code-highspeed: 1.90 in / 8.00 out / 0.38 cache hit | Kimi API (Moonshot): `kimi-k2.7-code`; Kimi API (Moonshot) high-speed: `kimi-k2.7-code-highspeed`; OpenRouter: `moonshotai/kimi-k2.7-code` | Hugging Face: https://huggingface.co/moonshotai/Kimi-K2.7-Code; Web app: https://www.kimi.com — Dedicated coding model; pairs with Kimi Code CLI. Release day not verified (HF 2026-06-11, OpenRouter 2026-06-12). Max output not verified.
  - Kimi API (Moonshot): `kimi-k2.7-code` — https://api.moonshot.ai/v1 (docs: https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart)
  - Kimi API (Moonshot) high-speed: `kimi-k2.7-code-highspeed` — https://api.moonshot.ai/v1 (docs: https://platform.kimi.ai/docs/models)
  - OpenRouter: `moonshotai/kimi-k2.7-code` — https://openrouter.ai/moonshotai/kimi-k2.7-code
  - Hugging Face — https://huggingface.co/moonshotai/Kimi-K2.7-Code
  - Web app — https://www.kimi.com
  - Pricing source: https://platform.kimi.ai/docs/pricing/chat
Notable capabilities:
  - Coding-specialized K2.6 derivative: Built on Kimi K2.6 (1T total / 32B active, MLA, 400M vision encoder) and tuned for long-horizon real-world coding. (https://huggingface.co/moonshotai/Kimi-K2.7-Code)
  - ~30% fewer thinking tokens than K2.6: Higher task success with about 30% lower thinking-token usage vs K2.6. (https://huggingface.co/moonshotai/Kimi-K2.7-Code)
  - High-speed tier: kimi-k2.7-code-highspeed outputs ~180 tok/s (up to ~260 tok/s on short context). (https://platform.kimi.ai/docs/models)

### Kimi K2.6
**Kimi K2.6** (Moonshot AI; current; multimodal; released 2026-04; open weights) | ctx 262,144 | $0.95 in / $4 out per 1M tokens (USD) | Kimi API (Moonshot): `kimi-k2.6`; Alibaba Cloud Model Studio: `kimi-k2.6`; OpenRouter: `moonshotai/kimi-k2.6` | Hugging Face: https://huggingface.co/moonshotai/Kimi-K2.6; Web app: https://www.kimi.com — Still offered on the API alongside K3 (only remaining non-coding K2-series model; kimi-k2.5 discontinued 2026-08-31). Release day not verified (HF 2026-04-14, OpenRouter 2026-04-20).
  - Kimi API (Moonshot): `kimi-k2.6` — https://api.moonshot.ai/v1 (docs: https://platform.kimi.ai/docs/guide/kimi-k2-6-quickstart)
  - Alibaba Cloud Model Studio: `kimi-k2.6` (docs: https://www.alibabacloud.com/help/en/model-studio/text-generation-model)
  - OpenRouter: `moonshotai/kimi-k2.6` — https://openrouter.ai/moonshotai/kimi-k2.6
  - Hugging Face — https://huggingface.co/moonshotai/Kimi-K2.6
  - Web app — https://www.kimi.com
  - Pricing source: https://platform.kimi.ai/docs/pricing/chat
Notable capabilities:
  - Native multimodal open agentic model: 1T total / 32B active MoE with 400M vision encoder; text, image and video input; thinking and non-thinking modes. (https://huggingface.co/moonshotai/Kimi-K2.6)
  - Swarm-based task orchestration: Marketed for proactive autonomous execution and agent-swarm orchestration plus coding-driven design. (https://huggingface.co/moonshotai/Kimi-K2.6)

### Kimi K2 Thinking
**Kimi K2 Thinking** (Moonshot AI; retired; reasoning-llm; released 2025-11; open weights) | ctx 262,144 | Kimi API (discontinued): `kimi-k2-thinking / kimi-k2-thinking-turbo`; OpenRouter: `moonshotai/kimi-k2-thinking` | Hugging Face: https://huggingface.co/moonshotai/Kimi-K2-Thinking — kimi-k2 series (incl. K2 Thinking, K2-0905, K2-0711) discontinued on the Kimi API on 2026-05-25; Moonshot recommends kimi-k3. Still available as open weights and via third parties. Pricing not verified.
  - Kimi API (discontinued): `kimi-k2-thinking / kimi-k2-thinking-turbo` — https://api.moonshot.ai/v1 (docs: https://platform.kimi.ai/docs/models)
  - OpenRouter: `moonshotai/kimi-k2-thinking` — https://openrouter.ai/moonshotai/kimi-k2-thinking
  - Hugging Face — https://huggingface.co/moonshotai/Kimi-K2-Thinking
Notable capabilities:
  - Long tool-call chains: Interleaves reasoning with function calls and stays coherent across 200-300 sequential tool calls (vs 30-50 for earlier models, per Moonshot). (https://huggingface.co/moonshotai/Kimi-K2-Thinking)
  - Native INT4 via quantization-aware training: QAT in post-training gives a lossless ~2x speed-up at INT4 on a 1T/32B-active MoE. (https://huggingface.co/moonshotai/Kimi-K2-Thinking)
  - Heavy mode: Parallel 8-trajectory rollout with reflective aggregation used for top benchmark results (HLE, BrowseComp). (https://huggingface.co/moonshotai/Kimi-K2-Thinking)


## Multimodal Art Projection (M-A-P)

### YuE2-3B
**YuE2-3B** (Multimodal Art Projection (M-A-P); current; music; released 2026-09-09; open weights) | Hugging Face: https://huggingface.co/m-a-p/YuE2-3B; GitHub (inference code, agent skill): https://github.com/multimodal-art-projection/YuE; Hugging Face (community GGUF): https://huggingface.co/audio-cpp/Yue2-3B-GGUF — Self-reported WildSongBench best-of-8 6.9632 vs Suno v5 6.8721. English + Mandarin lyrics. Model card states ~4B parameters. HF repos created 2026-09-09; exact public announcement day not verified. Predecessor YuE (2025-01-28, arXiv 2503.08638).
  - Hugging Face — https://huggingface.co/m-a-p/YuE2-3B
  - GitHub (inference code, agent skill) — https://github.com/multimodal-art-projection/YuE
  - Hugging Face (community GGUF) — https://huggingface.co/audio-cpp/Yue2-3B-GGUF
Notable capabilities:
  - Score-first song generation: Writes an editable melody-and-chord plan in ABC notation, then renders a full song with vocals and accompaniment (48 kHz stereo). (https://github.com/multimodal-art-projection/YuE)
  - Zero-shot covers and agentic editing: Covers from reference recordings (0.647 CLEWS mAP, self-reported) and conversational editing that turns musical feedback into score revisions. (https://huggingface.co/m-a-p/YuE2-3B)


## Nari Labs

### Nari Labs Dia2 (1B / 2B)
**Nari Labs Dia2 (1B / 2B)** (Nari Labs; current; audio/speech; released 2025-11-19; open weights) | Hugging Face: `nari-labs/Dia2-2B` | GitHub: https://github.com/nari-labs/dia2 — English only. Successor to Dia-1.6B (April 2025, github.com/nari-labs/dia). Release date 2025-11-19 from secondary sources (GitHub releases page).
  - Hugging Face: `nari-labs/Dia2-2B` — https://huggingface.co/nari-labs/Dia2-2B
  - GitHub — https://github.com/nari-labs/dia2
Notable capabilities:
  - Streaming multi-speaker dialogue TTS: Generates [S1]/[S2] dialogue and starts producing audio from the first few input tokens (no need for full text); conditions on audio prefixes for real-time conversation; up to ~2 min per generation (Mimi codec, 12.5 Hz); word-level timestamps. (https://huggingface.co/nari-labs/Dia2-2B)


## Neuphonic

### Neuphonic NeuTTS Air / NeuTTS Nano
**Neuphonic NeuTTS Air / NeuTTS Nano** (Neuphonic; current; audio/speech; released 2025-10-02; open weights) | Hugging Face: `neuphonic/neutts-nano-german` | GitHub: https://github.com/neuphonic/neutts — Release date from MarkTechPost coverage (2025-10-02). Nano license and exact Air HF repo id (neuphonic/neutts-air) not verified today.
  - GitHub — https://github.com/neuphonic/neutts
  - Hugging Face: `neuphonic/neutts-nano-german` — https://huggingface.co/neuphonic/neutts-nano-german
Notable capabilities:
  - On-device TTS with instant cloning: NeuTTS Air: 748M params (0.5B-class Qwen backbone + NeuCodec), real time from RTX 4090 down to Raspberry Pi, clones from ~3 s of audio, Perth watermark on every output; Nano: 229M total / 120M active for tighter edge devices. (https://www.marktechpost.com/2025/10/02/neuphonic-open-sources-neutts-air-a-748m-parameter-on-device-speech-language-model-with-instant-voice-cloning/)


## NVIDIA

### NVIDIA Nemotron 3.5 Lightning (30B-A3B)
**NVIDIA Nemotron 3.5 Lightning (30B-A3B)** (NVIDIA; current; llm; released 2026-08-11; open weights) | ctx 1,000,000 | $0.06 in / $0.16 out per 1M tokens (USD) on OpenRouter (also :free variant) | NVIDIA API (build.nvidia.com): `nvidia/nemotron-3.5-lightning-30b-a3b`; OpenRouter: `nvidia/nemotron-3.5-lightning` | Hugging Face: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 — Successor to Nemotron 3 Nano 30B-A3B (nvidia/nemotron-nano-3-30b-a3b on NIM). Knowledge cutoff = pre-training (Sep 2025); post-training to May 2026.
  - NVIDIA API (build.nvidia.com): `nvidia/nemotron-3.5-lightning-30b-a3b` — https://integrate.api.nvidia.com/v1/chat/completions (docs: https://build.nvidia.com)
  - OpenRouter: `nvidia/nemotron-3.5-lightning` — https://openrouter.ai/nvidia/nemotron-3.5-lightning
  - Hugging Face — https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
  - Pricing source: https://openrouter.ai/nvidia/nemotron-3.5-lightning
Notable capabilities:
  - Tiny-active MoE with 1M context: 30B total / 3B active hybrid Mamba-2 + attention MoE with up to 1M context (256K on a single H100). (https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)
  - Built for customization: Released with base checkpoint and NVFP4 builds (incl. speculative-decoding DSpark/DFlash variants); intended for fine-tuning and domain adaptation. (https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16)

### NVIDIA NemotronLabs VoiceChat 11B (and PersonaPlex-7B)
**NVIDIA NemotronLabs VoiceChat 11B (and PersonaPlex-7B)** (NVIDIA; current; audio/speech; released 2026-08-03; open weights) | Hugging Face: `nvidia/NVIDIA-NemotronLabs-VoiceChat-11B`; Hugging Face: `nvidia/personaplex-7b-v1` | arXiv: https://arxiv.org/abs/2609.21967 — English only. Requires datacenter GPU (A100/H100/H200/B100/B200 or RTX 6000). 'First' is NVIDIA's claim on the model card. HF card release date 2026-08-03; arXiv paper 2609.21967 (Sept 2026).
  - Hugging Face: `nvidia/NVIDIA-NemotronLabs-VoiceChat-11B` — https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
  - Hugging Face: `nvidia/personaplex-7b-v1` — https://huggingface.co/nvidia/personaplex-7b-v1
  - arXiv — https://arxiv.org/abs/2609.21967
Notable capabilities:
  - [FIRST] Open full-duplex speech model with tool calling: End-to-end (FastConformer encoder + Nemotron Nano v2 9B + TTS decoder, 11B total) full-duplex voice chat that calls tools mid-conversation; NVIDIA calls it the first open full-duplex model to support tool calling. BFCL-v3 (AU Harness) 56.1%, Full-Duplex-Bench v3 tool selection 82.5%. (https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B)
  - Natural turn-taking: ~450 ms turn-taking latency; #2 among open models on VoiceBench and Full-Duplex-Bench 1.0 (smooth turn-taking 0.82, interruption latency 480 ms). (https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B)
  - PersonaPlex: persona + voice prompted full duplex (found after launch): PersonaPlex-7B-v1 (2026-01-15), fine-tuned from Kyutai Moshiko, takes a voice prompt and a text persona/role prompt. (https://huggingface.co/nvidia/personaplex-7b-v1)

### NVIDIA Nemotron 3 Ultra (550B-A55B)
**NVIDIA Nemotron 3 Ultra (550B-A55B)** (NVIDIA; current; reasoning-llm; released 2026-06-04; open weights) | ctx 1,000,000 | $0.6 in / $2.4 out per 1M tokens (USD) on OpenRouter (262K context there); NVIDIA hosted pricing not verified | NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-ultra-550b-a55b`; OpenRouter: `nvidia/nemotron-3-ultra-550b-a55b` | Hugging Face: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 — Knowledge cutoff = pre-training data (Sep 2025); post-training data to May 2026. NVFP4 repo nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4.
  - NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-ultra-550b-a55b` — https://integrate.api.nvidia.com/v1/chat/completions (docs: https://build.nvidia.com)
  - OpenRouter: `nvidia/nemotron-3-ultra-550b-a55b` — https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b
  - Hugging Face — https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
  - Pricing source: https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b
Notable capabilities:
  - Hybrid Mamba-2 / LatentMoE at frontier scale: 550B total / 55B active; interleaved Mamba-2 and LatentMoE layers with select attention, plus multi-token prediction for faster generation. (https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16)
  - NVFP4 pretraining and weights: Pre-trained with an NVFP4 recipe; weights published in both BF16 and NVFP4 under the permissive OpenMDW-1.1 license. (https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16)
  - Reasoning on / off / medium: enable_thinking toggle in the chat template plus a medium-effort mode to cut reasoning tokens. (https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16)

### Cosmos 3 (Nano / Super)
**Cosmos 3 (Nano / Super)** (NVIDIA; current; world-model; released 2026-06-01; open weights) | Hugging Face: `nvidia/Cosmos3-Nano`; Hugging Face: `nvidia/Cosmos3-Super` | GitHub: https://github.com/nvidia-cosmos — Announced at GTC 2026-03-16 ('the first world foundation model unifying synthetic world generation, vision reasoning and action simulation' - NVIDIA claim); weights published 2026-05-31/06-01 (HF blog 'The First Open Omni-model for Physical AI Reasoning and Action'). Sizes: Nano 16B, Super 64B. Linux + Ampere/Hopper/Blackwell GPUs, BF16. Technical report dated 2026-06-22.
  - Hugging Face: `nvidia/Cosmos3-Nano` — https://huggingface.co/nvidia/Cosmos3-Nano
  - Hugging Face: `nvidia/Cosmos3-Super` — https://huggingface.co/nvidia/Cosmos3-Super
  - GitHub — https://github.com/nvidia-cosmos (docs: https://research.nvidia.com/labs/cosmos-lab/cosmos3/)
Notable capabilities:
  - [FIRST] Unified omni world model (generation + reasoning + action): One Mixture-of-Transformers model (autoregressive + diffusion) replaces separate Cosmos Predict, Transfer, Reason and Policy models: world generation, physical reasoning, forward/inverse dynamics and action/policy generation. (https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world)
  - Open omnimodal I/O: Inputs text, images, short video, audio and action trajectories (16-400 frames); outputs text, images, video (5-400 frames), 48 kHz stereo audio and actions (JSON). (https://huggingface.co/nvidia/Cosmos3-Nano)
  - Leaderboard results (found after launch): NVIDIA cites best open text-to-image and image-to-video models on Artificial Analysis and best policy model on RoboArena. (https://www.nvidia.com/en-us/ai/cosmos/)

### NVIDIA Nemotron 3 Nano Omni (30B-A3B Reasoning)
**NVIDIA Nemotron 3 Nano Omni (30B-A3B Reasoning)** (NVIDIA; current; multimodal; released 2026-04-28; open weights) | ctx 256,000 | NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning`; OpenRouter: `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` | Hugging Face: https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 — Also FP8/NVFP4 repos. Only free OpenRouter variant seen; paid pricing not verified. Knowledge cutoff not published.
  - NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning` — https://integrate.api.nvidia.com/v1/chat/completions (docs: https://build.nvidia.com)
  - OpenRouter: `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free` — https://openrouter.ai/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
  - Hugging Face — https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
Notable capabilities:
  - Open omni-modal reasoning (video + audio + image): Single 3B-active open model reasoning over video (up to ~2 min), audio, images and text with chain-of-thought on by default. (https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16)
  - ASR with word timestamps, OCR, GUI automation: Targets transcription with word-level timestamps, document intelligence/OCR and GUI agent workflows. (https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16)

### Isaac GR00T N1.7
**Isaac GR00T N1.7** (NVIDIA; current; robotics; released 2026-04-17; open weights) | Hugging Face: `nvidia/GR00T-N1.7-3B` | GitHub: https://github.com/NVIDIA/Isaac-GR00T — Early access with commercial licensing announced at GTC 2026-03-16; open release/HF blog 2026-04-17. Post-trained checkpoints: GR00T-N1.7-LIBERO, -DROID, -SimplerEnv-Bridge, -SimplerEnv-Fractal, GR00T-H-N1.7 (surgical-robotics variant, uploaded to HF 2026-05-30: 3B, post-trained on 601 h / ~63.9k episodes of real surgical tasks from the Open-H-Embodiment dataset across 7 platforms incl. dVRK, CMR Versius, KUKA LBR iiwa; NVIDIA Open Model License; R&D only, not for clinical use; follows the original GR00T-H announced at GTC 2026-03-16). Backbone nvidia/Cosmos-Reason2-2B is gated (accept license on HF). Validated on Unitree G1, YAM bimanual, AGIBot Genie 1. Fine-tuning: 40 GB+ GPUs recommended.
  - Hugging Face: `nvidia/GR00T-N1.7-3B` — https://huggingface.co/nvidia/GR00T-N1.7-3B (docs: https://huggingface.co/blog/nvidia/gr00t-n1-7)
  - GitHub — https://github.com/NVIDIA/Isaac-GR00T
  - Hugging Face LeRobot integration (docs: https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/)
Notable capabilities:
  - Human egocentric video pretraining: Pretrained on 20,854 hours of human egocentric video (EgoScale) across 20+ task categories, on top of robot data. (https://huggingface.co/blog/nvidia/gr00t-n1-7)
  - [FIRST] Scaling law for robot dexterity: NVIDIA reports the 'first-ever scaling law for robot dexterity': more human video predictably improves 22-DoF hand performance without mass teleoperation. (https://huggingface.co/blog/nvidia/gr00t-n1-7)
  - Reasoning VLA on a Cosmos backbone: 3B 'Action Cascade' model: Cosmos-Reason2-2B VLM plus 32-layer diffusion transformer; relative end-effector action space; runs on one 16 GB+ GPU including Jetson Thor/Orin and DGX Spark. (https://github.com/NVIDIA/Isaac-GR00T)

### NVIDIA Nemotron 3 Super (120B-A12B)
**NVIDIA Nemotron 3 Super (120B-A12B)** (NVIDIA; current; reasoning-llm; released 2026-03-11; open weights) | ctx 1,000,000 | $0.08 in / $0.45 out per 1M tokens (USD) on OpenRouter; NVIDIA hosted pricing not verified | NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-super-120b-a12b`; AWS Bedrock: `nvidia.nemotron-super-3-120b`; OpenRouter: `nvidia/nemotron-3-super-120b-a12b` | Hugging Face: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 — Knowledge cutoff = pre-training (Jun 2025); post-training to Feb 2026. Also FP8/NVFP4 repos. Free tier on OpenRouter (:free).
  - NVIDIA API (build.nvidia.com): `nvidia/nemotron-3-super-120b-a12b` — https://integrate.api.nvidia.com/v1/chat/completions (docs: https://build.nvidia.com)
  - AWS Bedrock: `nvidia.nemotron-super-3-120b` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html)
  - OpenRouter: `nvidia/nemotron-3-super-120b-a12b` — https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b
  - Hugging Face — https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
  - Pricing source: https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b
Notable capabilities:
  - Efficient hybrid LatentMoE for agents: 120B total / 12B active hybrid Mamba-2 + MoE + attention, built for high-volume agentic workloads with up to 1M context (256K default). (https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16)
  - Managed on AWS Bedrock (found after launch): One of the few NVIDIA open models offered as a serverless Bedrock model (nvidia.nemotron-super-3-120b). (https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html)

### Cosmos Reason 2
**Cosmos Reason 2** (NVIDIA; current; multimodal; released 2025-12-19; open weights) | ctx 256,000 | Hugging Face: `nvidia/Cosmos-Reason2-8B`; Hugging Face: `nvidia/Cosmos-Reason2-2B` — Based on Qwen3-VL (8B variant from Qwen3-VL-8B-Instruct, 8.7B params, 32 GB+ GPU). Initial release 2025-12-19, updated 2026-03-10; promoted at CES 2026. Up to 256K input tokens. Its role is folded into Cosmos 3 for new projects.
  - Hugging Face: `nvidia/Cosmos-Reason2-8B` — https://huggingface.co/nvidia/Cosmos-Reason2-8B
  - Hugging Face: `nvidia/Cosmos-Reason2-2B` — https://huggingface.co/nvidia/Cosmos-Reason2-2B
Notable capabilities:
  - Physical-AI reasoning VLM: Spatio-temporal video reasoning, 2D/3D point and box localization, robot planning; 8B beats base Qwen3-VL-8B on robotics (56.90 vs 53.08) and self-driving (67.85 vs 46.38) evals per model card. (https://huggingface.co/nvidia/Cosmos-Reason2-8B)
  - Backbone for GR00T N1.7 (found after launch): Cosmos-Reason2-2B is the VLM backbone of Isaac GR00T N1.7. (https://github.com/NVIDIA/Isaac-GR00T)

### NVIDIA MagpieTTS Multilingual 357M
**NVIDIA MagpieTTS Multilingual 357M** (NVIDIA; current; audio/speech; released 2025-12-11; open weights) | Hugging Face: `nvidia/magpie_tts_multilingual_357m` | Hugging Face collection: https://huggingface.co/collections/nvidia/nemotron-speech — Versions: v2512 (HF repo created 2025-12-11), v2602 (Mar 2026), v2607 (2026-07-21); repo last updated 2026-09-09. Zero-shot voice cloning was deliberately removed 'for security reasons'. Part of the Nemotron Speech collection with Parakeet ASR, PersonaPlex and NemotronLabs-VoiceChat.
  - Hugging Face: `nvidia/magpie_tts_multilingual_357m` — https://huggingface.co/nvidia/magpie_tts_multilingual_357m
  - Hugging Face collection — https://huggingface.co/collections/nvidia/nemotron-speech
Notable capabilities:
  - Small open multilingual TTS for commercial use: ~357-364M-parameter transformer encoder-decoder predicting multi-codebook audio codec tokens; 12 languages (ar, zh, en, fr, de, hi, it, ja, ko, pt, es, vi); 5 built-in English voices; CER 0.34-3.17% across languages per model card; trained on ~54,300 h. (https://huggingface.co/nvidia/magpie_tts_multilingual_357m)

### NVIDIA Parakeet / Canary / Nemotron Speech ASR (open)
**NVIDIA Parakeet / Canary / Nemotron Speech ASR (open)** (NVIDIA; current; audio/speech; released 2025-08-14; open weights) | Hugging Face: `nvidia/parakeet-tdt-0.6b-v3`; Hugging Face: `nvidia/canary-qwen-2.5b`; Hugging Face: `nvidia/parakeet-unified-en-0.6b`; Hugging Face: `nvidia/nemotron-speech-streaming-en-0.6b`; Hugging Face: `nvidia/nemotron-3.5-asr-streaming-0.6b` | NVIDIA NIM / build.nvidia.com: https://build.nvidia.com — One file for NVIDIA's open ASR family. Also canary-1b-v2 (European ASR + translation) and parakeet-tdt-0.6b-v2 (English, NIM). Nemotron 3.5 ASR HF card shows a garbled date; June 2026 per NVIDIA/press. NVIDIA's open TTS: magpie_tts_multilingual_357m. Full-duplex model: see nemotron-voicechat.
  - Hugging Face: `nvidia/parakeet-tdt-0.6b-v3` — https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
  - Hugging Face: `nvidia/canary-qwen-2.5b` — https://huggingface.co/nvidia/canary-qwen-2.5b
  - Hugging Face: `nvidia/parakeet-unified-en-0.6b` — https://huggingface.co/nvidia/parakeet-unified-en-0.6b
  - Hugging Face: `nvidia/nemotron-speech-streaming-en-0.6b` — https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b
  - Hugging Face: `nvidia/nemotron-3.5-asr-streaming-0.6b` — https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b
  - NVIDIA NIM / build.nvidia.com — https://build.nvidia.com
Notable capabilities:
  - Parakeet TDT 0.6B v3: 25 European languages, very high throughput: 600M FastConformer-TDT with auto language ID, punctuation, word timestamps, up to 24 min (3 h with local attention); 6.34% avg WER on Open ASR Leaderboard; trained on the Granary dataset. (https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)
  - Canary-Qwen-2.5B speech-augmented LLM: FastConformer encoder + Qwen LLM (SALM); 5.63% mean WER topped the HF Open ASR Leaderboard at release (2025-07-17); can summarize/answer questions about the transcript. English, max 40 s clips. (https://huggingface.co/nvidia/canary-qwen-2.5b)
  - Cache-aware streaming ASR, 80-1120 ms chunks (found after launch): Nemotron Speech Streaming EN 0.6B (Jan/Mar 2026) and Nemotron 3.5 ASR Streaming 0.6B (June 2026, 40 language-locales) switch latency at inference without retraining; Parakeet-unified-en-0.6B (2026-04-07) does both offline (5.91% WER) and streaming down to 160 ms. (https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b)

### Isaac GR00T N2
**Isaac GR00T N2** (NVIDIA; preview; robotics; released 2026-03-16) | Not yet available (NVIDIA says end of 2026): https://developer.nvidia.com/isaac/gr00t — Previewed in Jensen Huang's GTC keynote 2026-03-16; 'released' = preview date. No weights, API or HF repo found as of 2026-09-29. Modalities assumed from the GR00T line; confirm at release.
  - Not yet available (NVIDIA says end of 2026) — https://developer.nvidia.com/isaac/gr00t
Notable capabilities:
  - World action model (DreamZero): Predicts how the scene will evolve (future latent states) before generating the action sequence; succeeds at new tasks in new environments more than twice as often as leading VLAs (NVIDIA). (https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world)
  - Top of generalist-policy leaderboards: NVIDIA says it ranks No. 1 on MolmoSpaces and RoboArena for generalist robot policies (as of GTC, March 2026). (https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world)

### Cosmos Predict 2.5 / Transfer 2.5
**Cosmos Predict 2.5 / Transfer 2.5** (NVIDIA; legacy; world-model; released 2025-10-06; open weights) | Hugging Face: `nvidia/Cosmos-Predict2.5-2B`; Hugging Face: `nvidia/Cosmos-Predict2.5-14B`; Hugging Face: `nvidia/Cosmos-Transfer2.5-2B` | GitHub: https://github.com/nvidia-cosmos/cosmos-transfer2.5 — Predict 2.5-2B released 2025-10-06 (per model card); needs ~32.5 GB VRAM. Consolidated into Cosmos 3 (June 2026) but still downloadable.
  - Hugging Face: `nvidia/Cosmos-Predict2.5-2B` — https://huggingface.co/nvidia/Cosmos-Predict2.5-2B
  - Hugging Face: `nvidia/Cosmos-Predict2.5-14B` — https://huggingface.co/nvidia/Cosmos-Predict2.5-14B
  - Hugging Face: `nvidia/Cosmos-Transfer2.5-2B` — https://huggingface.co/nvidia/Cosmos-Transfer2.5-2B
  - GitHub — https://github.com/nvidia-cosmos/cosmos-transfer2.5
Notable capabilities:
  - Unified Text2World / Image2World / Video2World: Single diffusion transformer for physics-aware video world generation (720p, 16 fps, ~5 s clips) for robotics and AV synthetic data. (https://huggingface.co/nvidia/Cosmos-Predict2.5-2B)
  - Multi-control world-to-world transfer: Transfer 2.5 generates world simulations conditioned on spatial controls (depth, segmentation, edges etc.) on top of Predict 2.5. (https://github.com/nvidia-cosmos/cosmos-transfer2.5)

### Isaac GR00T N1 / N1.5 / N1.6
**Isaac GR00T N1 / N1.5 / N1.6** (NVIDIA; legacy; robotics; released 2025-03-18; open weights) | Hugging Face: `nvidia/GR00T-N1-2B`; Hugging Face: `nvidia/GR00T-N1.5-3B`; Hugging Face: `nvidia/GR00T-N1.6-3B` | GitHub (branches n1d5, n1d6): https://github.com/NVIDIA/Isaac-GR00T — N1 (2B) announced 2025-03-18; N1.5 (3B) mid-2025; N1.6 (3B) later in 2025 - exact N1.5/N1.6 dates not re-verified. Superseded by GR00T N1.7 (2026). 'first' is NVIDIA's claim (open weights for a humanoid-specific generalist model; earlier open VLAs such as OpenVLA/Octo targeted arms).
  - Hugging Face: `nvidia/GR00T-N1-2B` — https://huggingface.co/nvidia/GR00T-N1-2B
  - Hugging Face: `nvidia/GR00T-N1.5-3B` — https://huggingface.co/nvidia/GR00T-N1.5-3B
  - Hugging Face: `nvidia/GR00T-N1.6-3B` — https://huggingface.co/nvidia/GR00T-N1.6-3B
  - GitHub (branches n1d5, n1d6) — https://github.com/NVIDIA/Isaac-GR00T
Notable capabilities:
  - [FIRST] Open humanoid robot foundation model: Announced at GTC 2025 as 'the world's first open humanoid robot foundation model': a dual-system VLA (VLM 'System 2' + diffusion-transformer 'System 1') for cross-embodiment humanoid control, customizable with synthetic data. (https://nvidianews.nvidia.com/news/nvidia-isaac-gr00t-n1-open-humanoid-robot-foundation-model-simulation-frameworks)


## OpenAI

### GPT-6 Luna
**GPT-6 Luna** (OpenAI; current; reasoning-llm; released 2026-09-22) | ctx 1,050,000 | $0.1 in / $0.5 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-6-luna`; Azure OpenAI (Microsoft Foundry): `gpt-6-luna`; OpenRouter: `openai/gpt-6-luna` | Web app: https://chatgpt.com — Most efficient GPT-6 model for focused, high-volume tasks; successor to GPT-5.6 Luna (the mini/nano tier). OpenRouter also lists openai/gpt-6-luna-pro.
  - OpenAI API: `gpt-6-luna` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-6-luna)
  - Azure OpenAI (Microsoft Foundry): `gpt-6-luna` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-6-luna` — https://openrouter.ai/openai/gpt-6-luna
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - 1M context at $0.10/M: Cheapest OpenAI reasoning model with the full 1.05M context window and 128K output. (https://developers.openai.com/api/docs/models/gpt-6-luna)
  - Agentic tools on the budget tier: Supports computer use, hosted shell, MCP and tool search like the larger models. (https://developers.openai.com/api/docs/models/gpt-6-luna)
  - Free-tier ChatGPT model: Rolled out to ChatGPT free users and the desktop app at launch. (https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/)

### GPT-6 Sol
**GPT-6 Sol** (OpenAI; current; reasoning-llm; released 2026-09-22) | ctx 1,050,000 | $2 in / $10 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-6-sol`; Azure OpenAI (Microsoft Foundry): `gpt-6-sol`; OpenRouter: `openai/gpt-6-sol` | Web app: https://chatgpt.com — Mid-tier GPT-6 model for complex coding and agentic workflows; successor to GPT-5.6 Sol. Reasoning effort none..max. OpenRouter also lists openai/gpt-6-sol-pro (reasoning.mode pro).
  - OpenAI API: `gpt-6-sol` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-6-sol)
  - Azure OpenAI (Microsoft Foundry): `gpt-6-sol` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-6-sol` — https://openrouter.ai/openai/gpt-6-sol
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Astra-level reliability at lower cost: OpenAI claims about half as many mistakes as GPT-5.6 Sol at half its API price. (https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/)
  - Full hosted tool suite: Web/file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search. (https://developers.openai.com/api/docs/models/gpt-6-sol)
  - Image-input bug fix (found after launch): Sep 25 2026 fix for an image-encoding bug that degraded image understanding at launch. (https://developers.openai.com/api/docs/changelog)

### GPT Image 2.5 Flare
**GPT Image 2.5 Flare** (OpenAI; current; image-gen; released 2026-09-08) | OpenAI API: `gpt-image-2.5-flare`; Azure OpenAI (Microsoft Foundry): `gpt-image-2.5-flare`; ElevenLabs Image & Video API: `gpt-image-2.5-flare` | Web app: https://chatgpt.com — Snapshot gpt-image-2.5-flare-2026-09-08. Same token rates as Sunburst and gpt-image-2. OpenRouter id not verified.
  - OpenAI API: `gpt-image-2.5-flare` — https://api.openai.com/v1/images/generations (docs: https://developers.openai.com/api/docs/models/gpt-image-2.5-flare)
  - Azure OpenAI (Microsoft Foundry): `gpt-image-2.5-flare` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Web app — https://chatgpt.com
  - ElevenLabs Image & Video API: `gpt-image-2.5-flare` — /image/create (docs: https://elevenlabs.io/docs/overview/capabilities/image-video)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Fast everyday image generation: Fastest high-quality OpenAI image model; quality low/medium/high/xhigh/max/auto. (https://developers.openai.com/api/docs/models/gpt-image-2.5-flare)
  - Inpainting: Editing with masks via v1/images/edits. (https://developers.openai.com/api/docs/models/gpt-image-2.5-flare)

### GPT Image 2.5 Sunburst
**GPT Image 2.5 Sunburst** (OpenAI; current; image-gen; released 2026-09-08) | OpenAI API: `gpt-image-2.5-sunburst`; Azure OpenAI (Microsoft Foundry): `gpt-image-2.5-sunburst`; ElevenLabs Image & Video API: `gpt-image-2.5-sunburst` | Web app: https://chatgpt.com — Snapshot gpt-image-2.5-sunburst-2026-09-08. OpenRouter id not verified.
  - OpenAI API: `gpt-image-2.5-sunburst` — https://api.openai.com/v1/images/generations (docs: https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst)
  - Azure OpenAI (Microsoft Foundry): `gpt-image-2.5-sunburst` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Web app — https://chatgpt.com
  - ElevenLabs Image & Video API: `gpt-image-2.5-sunburst` — /image/create (docs: https://elevenlabs.io/docs/overview/capabilities/image-video)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Most capable OpenAI image model: Top-quality generation and editing with inpainting via images/generations and images/edits. (https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst)
  - Replacement for gpt-image-1.5/1-mini (found after launch): Named successor for image models shutting down Dec 1 2026. (https://developers.openai.com/api/docs/deprecations)

### GPT-6 Astra
**GPT-6 Astra** (OpenAI; current; reasoning-llm; released 2026-09-03) | ctx 1,050,000 | $10 in / $50 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-6-astra`; Azure OpenAI (Microsoft Foundry): `gpt-6-astra`; OpenRouter: `openai/gpt-6-astra` | Web app: https://chatgpt.com — OpenAI flagship ("most capable model, built for the hardest end-to-end work"). API changelog: Sep 3 2026 (limited preview Sep 3, public Sep 4). Single snapshot gpt-6-astra. OpenRouter also lists openai/gpt-6-astra-pro = same model with reasoning.mode pro. Endpoints: Chat Completions, Responses, Batch.
  - OpenAI API: `gpt-6-astra` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-6-astra)
  - Azure OpenAI (Microsoft Foundry): `gpt-6-astra` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-6-astra` — https://openrouter.ai/openai/gpt-6-astra
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Max reasoning effort: reasoning.effort adds a new "max" level above xhigh (low/medium/high/xhigh/max). (https://developers.openai.com/api/docs/models/gpt-6-astra)
  - 1M-token context: 1.05M context window (922K max input) with 128K output on the flagship. (https://developers.openai.com/api/docs/models/gpt-6-astra)
  - Restricted cyber behaviour: Released as a restricted version that rejects certain cybersecurity prompts; separate Cyber/Daybreak models exist for that domain. (https://en.wikipedia.org/wiki/GPT-6_Astra)
  - Recurrent-depth reasoning (found after launch): Reported new "recurrent depth" technique that obscures some of the reasoning, raising monitorability concerns among safety researchers. (https://en.wikipedia.org/wiki/GPT-6_Astra)

### GPT-Live-Transcribe
**GPT-Live-Transcribe** (OpenAI; current; audio/speech; released 2026-07-28) | OpenAI API: `gpt-live-transcribe` — Released with gpt-transcribe (file transcription, $0.0045/min) on 2026-07-28 per the changelog. Languages and latency figures not published on the docs page.
  - OpenAI API: `gpt-live-transcribe` — v1/realtime/transcription_sessions (Realtime transcription session over WebSocket/WebRTC) (docs: https://developers.openai.com/api/docs/models/gpt-live-transcribe)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Low-latency streaming transcription with context hints: Streams transcript deltas with tunable latency and accepts unstructured context, keyword hints and multiple language hints. (https://developers.openai.com/api/docs/models/gpt-live-transcribe)
  - Recommended replacement for Whisper streaming use (found after launch): Named (with gpt-transcribe) as the replacement for whisper-1 and gpt-4o-(mini-)transcribe(-diarize), which shut down 2027-02-26. (https://developers.openai.com/api/docs/deprecations)

### GPT-Transcribe
**GPT-Transcribe** (OpenAI; current; audio/speech; released 2026-07-28) | OpenAI API: `gpt-transcribe`; Azure OpenAI (Microsoft Foundry): `gpt-transcribe` — File and Realtime transcription. Streaming sibling gpt-live-transcribe ($0.017/min). Cheaper than whisper-1 ($0.006/min).
  - OpenAI API: `gpt-transcribe` — https://api.openai.com/v1/audio/transcriptions (docs: https://developers.openai.com/api/docs/models/gpt-transcribe)
  - Azure OpenAI (Microsoft Foundry): `gpt-transcribe` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Context-guided transcription: Accepts unstructured context, keyword hints and multiple language hints for domain terms. (https://developers.openai.com/api/docs/models/gpt-transcribe)
  - Whisper successor (found after launch): Replacement for whisper-1 and gpt-4o-(mini-)transcribe (shutdown Feb 26 2027). (https://developers.openai.com/api/docs/deprecations)

### GPT-5.6 Terra
**GPT-5.6 Terra** (OpenAI; current; reasoning-llm; released 2026-07-09) | ctx 1,050,000 | $2 in / $12 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.6-terra`; Azure OpenAI (Microsoft Foundry): `gpt-5.6-terra`; OpenRouter: `openai/gpt-5.6-terra` | Web app: https://chatgpt.com — Balanced GPT-5.6 model; no GPT-6 Terra counterpart as of 2026-09-29. Single snapshot gpt-5.6-terra.
  - OpenAI API: `gpt-5.6-terra` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.6-terra)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.6-terra` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.6-terra` — https://openrouter.ai/openai/gpt-5.6-terra
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Balanced tier: Mid tier at $2/$12 per 1M, well below GPT-5.5 ($5/$30), with max reasoning effort. (https://developers.openai.com/api/docs/pricing)
  - Official migration target (found after launch): Named replacement for many deprecated legacy snapshots (gpt-3.5, gpt-4 variants, o-series). (https://developers.openai.com/api/docs/deprecations)

### GPT-Live 1
**GPT-Live 1** (OpenAI; current; audio/speech; released 2026-07-08) | OpenAI API: `gpt-live-1` | Web app: https://chatgpt.com — Launched in ChatGPT 2026-07-08 (GPT-Live-1 for Go/Plus/Pro, GPT-Live-1 mini default for Free); ChatGPT desktop (macOS/Windows) ~2026-07-23; API GA 2026-09-10 per changelog (earlier preview around 2026-07-31). gpt-live-1-mini is ChatGPT-only: not in the API models catalog and developers.openai.com/api/docs/models/gpt-live-1-mini returns 404 (checked 2026-09-29). ChatGPT Voice limits (Unite.AI, 2026-09-23): Free limited mini, Go 3 h mini, Plus 3 h GPT-Live-1, Pro $100 15 h, Pro $200 unlimited; Enterprise/Edu 1.25 credits/min or $0.05/min. Since 2026-09-23 Voice can use plugins/connected apps and runs inside ChatGPT Work. Knowledge cutoff 2025-07-31. Concurrency 25-500 sessions by tier. No image/video input. Not listed on Azure or OpenRouter. OpenAI's launch post returned 403 to our fetcher; ChatGPT facts from TechCrunch.
  - OpenAI API: `gpt-live-1` — https://api.openai.com/v1/live/sessions (docs: https://developers.openai.com/api/docs/models/gpt-live-1)
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/models/gpt-live-1
Notable capabilities:
  - Full-duplex voice: Listens and speaks at the same time, delegating reasoning and tool use to a backend agent model. (https://developers.openai.com/api/docs/models/gpt-live-1)
  - New Live API: Served on a dedicated v1/live/sessions endpoint rather than Realtime. (https://developers.openai.com/api/docs/models/gpt-live-1)
  - Replaced turn-based Advanced Voice Mode in ChatGPT: Since 2026-07-08 GPT-Live-1 (paid tiers) and GPT-Live-1 mini (default, all users) power ChatGPT Voice, with backchannels ('mhmm') and background hand-off of hard questions to GPT-5.5. (https://techcrunch.com/2026/07/08/openai-releases-new-voice-models-for-more-natural-live-conversations/)

### GPT-Realtime-2.1
**GPT-Realtime-2.1** (OpenAI; current; audio/speech; released 2026-07-06) | ctx 128,000 | OpenAI API: `gpt-realtime-2.1`; Azure OpenAI (Microsoft Foundry): `gpt-realtime-2.1` — Realtime API only. Successor to gpt-realtime-2 (2026-05-07, same prices, see gpt-realtime-2.md). Mini variant gpt-realtime-2.1-mini (audio $10/$20, text $0.60/$2.40). Replaces gpt-realtime / gpt-4o-realtime (shutdown Jan 20 2027). Azure version 2026-07-07.
  - OpenAI API: `gpt-realtime-2.1` — wss://api.openai.com/v1/realtime (docs: https://developers.openai.com/api/docs/models/gpt-realtime-2.1)
  - Azure OpenAI (Microsoft Foundry): `gpt-realtime-2.1` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Reasoning in realtime voice: Configurable reasoning effort in a speech-to-speech model (at a latency cost). (https://developers.openai.com/api/docs/models/gpt-realtime-2.1)
  - Robust turn-taking: Improved alphanumeric recognition, silence/noise handling and interruption behavior. (https://developers.openai.com/api/docs/models/gpt-realtime-2.1)

### GPT-Realtime-2
**GPT-Realtime-2** (OpenAI; current; audio/speech; released 2026-05-07) | ctx 128,000 | OpenAI API: `gpt-realtime-2` — Launched 2026-05-07 with gpt-realtime-translate and gpt-realtime-whisper (changelog). Superseded two months later by gpt-realtime-2.1 (2026-07-06) at identical prices, but still listed and not deprecated. Realtime endpoint only; function calling and prompt caching. Official launch post (openai.com) returned 403 to our fetcher, so benchmark claims were not read directly; secondary sources quote OpenAI: +15.2% Big Bench Audio vs gpt-realtime-1.5 (high effort), +13.8% Audio MultiChallenge instruction following (xhigh); one blog reports 96.6% absolute Big Bench Audio at xhigh (unconfirmed).
  - OpenAI API: `gpt-realtime-2` — wss://api.openai.com/v1/realtime (docs: https://developers.openai.com/api/docs/models/gpt-realtime-2)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Reasoning speech-to-speech model: First OpenAI realtime voice model with configurable reasoning effort (press: 'GPT-5-class' reasoning); higher effort adds latency and tokens. (https://developers.openai.com/api/docs/models/gpt-realtime-2)
  - 128K-token realtime context: Context grew from 32K (gpt-realtime-1.5) to 128K tokens, with 32K max output, for long voice-agent sessions. (https://www.ghacks.net/2026/05/11/openai-releases-three-new-realtime-voice-models-for-the-api-with-gpt-5-class-reasoning/)

### GPT-Realtime-Translate
**GPT-Realtime-Translate** (OpenAI; current; audio/speech; released 2026-05-07) | ctx 16,000 | OpenAI API: `gpt-realtime-translate` — Language counts (70+ in / 13 out) come from press coverage of the launch post; the docs page does not list languages. Latency not specified. Google's comparable model is gemini-3.5-live-translate-preview (June 2026).
  - OpenAI API: `gpt-realtime-translate` — https://api.openai.com/v1/realtime/translations (docs: https://developers.openai.com/api/docs/models/gpt-realtime-translate)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Streaming speech-to-speech translation: Simultaneous interpretation from 70+ input languages into 13 output languages, emitting translated audio plus transcript deltas while the speaker is still talking. (https://www.ghacks.net/2026/05/11/openai-releases-three-new-realtime-voice-models-for-the-api-with-gpt-5-class-reasoning/)
  - Dedicated translation endpoint: Served only on v1/realtime/translations (not the general Realtime or Chat endpoints). (https://developers.openai.com/api/docs/models/gpt-realtime-translate)

### GPT-Rosalind
**GPT-Rosalind** (OpenAI; current; reasoning-llm; released 2026-04-17) | $5 in / $25 out per 1M tokens (USD); billing starts 2026-10-05 | OpenAI API (trusted access only): `gpt-rosalind-research` | ChatGPT / Codex (eligible organisations): https://openai.com/gpt-rosalind/ — Research preview 17 Apr 2026; rebuilt on GPT-5.5 on 3 June 2026 (OpenAI says 31% fewer tokens than GPT-5.5); out of preview globally 11 Sept 2026. Context window and max output not published. Pricing per OpenAI's pricing page 'Life Sciences' section, as quoted by TokenCost and the Portkey model registry (PR #953); not read directly on openai.com (403). Free Codex Life Sciences plugin connects any model to 50+ scientific tools.
  - OpenAI API (trusted access only): `gpt-rosalind-research` — https://api.openai.com/v1/chat/completions (docs: https://help.openai.com/en/articles/20001193-introducing-gpt-rosalind-for-life-sciences-research)
  - ChatGPT / Codex (eligible organisations) — https://openai.com/gpt-rosalind/
  - Pricing source: https://tokencost.app/blog/gpt-rosalind-pricing-billing-october-5
Notable capabilities:
  - Life-sciences specialist reasoning: Tuned for genomics, protein and sequence analysis, medicinal chemistry, literature synthesis, wet-lab troubleshooting and experiment planning; OpenAI reports BixBench pass@1 0.751 at launch and LabWorkBench 63.2% (vs GPT-5.5 55.8%) after the June update. (https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind/)
  - Trusted-access dual-use deployment: Callable only by vetted organisations with an approved research deployment; a Rosalind Biodefense programme extends access to US government and allied public-health partners. (https://www.rdworldonline.com/openai-launches-rosalind-biodefense-offers-federal-agencies-early-access-to-its-life-sciences-model/)

### GPT-Audio-1.5 (and gpt-audio / gpt-audio-mini)
**GPT-Audio-1.5 (and gpt-audio / gpt-audio-mini)** (OpenAI; current; audio/speech; released 2026-02-23) | ctx 128,000 | OpenAI API: `gpt-audio-1.5`; OpenAI API: `gpt-audio-mini` — gpt-audio-1.5 released 2026-02-23 with gpt-realtime-1.5. Older gpt-audio (2025) and gpt-audio-mini (2025-10-06) were deprecated 2026-07-20 with shutdown 2027-01-20 (replacement gpt-audio-1.5); gpt-4o-audio-preview was shut down 2026-05-12. Chat Completions only (not Responses).
  - OpenAI API: `gpt-audio-1.5` — https://api.openai.com/v1/chat/completions (docs: https://developers.openai.com/api/docs/models/gpt-audio-1.5)
  - OpenAI API: `gpt-audio-mini` — https://api.openai.com/v1/chat/completions (docs: https://developers.openai.com/api/docs/models/gpt-audio-mini)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Audio in / audio out over Chat Completions: Non-realtime REST alternative to the Realtime API: send audio and receive spoken audio plus text in one Chat Completions call, with streaming and function calling. (https://developers.openai.com/api/docs/models/gpt-audio-1.5)

### GPT-5.3-Codex
**GPT-5.3-Codex** (OpenAI; current; code; released 2026-02-05) | ctx 400,000 | $1.75 in / $14 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.3-codex`; Azure OpenAI (Microsoft Foundry): `gpt-5.3-codex`; OpenRouter: `openai/gpt-5.3-codex` | Codex (ChatGPT): https://chatgpt.com/codex — Latest codex-specific API id on the pricing page. Released in Codex Feb 5 2026; API access followed later (Azure version 2026-02-24). GPT-6 Sol is now positioned for coding.
  - OpenAI API: `gpt-5.3-codex` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.3-codex)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.3-codex` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.3-codex` — https://openrouter.ai/openai/gpt-5.3-codex
  - Codex (ChatGPT) — https://chatgpt.com/codex
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Agentic coding specialist: Codex-tuned GPT-5.3 for long-running software engineering (Codex app/CLI/IDE and API). (https://developers.openai.com/api/docs/models/gpt-5.3-codex)
  - Responses-only: Available only through the Responses API; effort low/medium/high/xhigh. (https://developers.openai.com/api/docs/models/gpt-5.3-codex)

### gpt-oss-120b
**gpt-oss-120b** (OpenAI; current; reasoning-llm; released 2025-08-05; open weights) | ctx 131,072 | OpenAI API (docs): `gpt-oss-120b`; Azure OpenAI (Microsoft Foundry): `gpt-oss-120b`; OpenRouter: `openai/gpt-oss-120b` | Hugging Face: https://huggingface.co/openai/gpt-oss-120b — Open weights (Apache 2.0). OpenRouter from ~$0.04/$0.17 per 1M (provider-dependent). No first-party OpenAI pricing listed.
  - Hugging Face — https://huggingface.co/openai/gpt-oss-120b
  - OpenAI API (docs): `gpt-oss-120b` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-oss-120b)
  - Azure OpenAI (Microsoft Foundry): `gpt-oss-120b` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-oss-120b` — https://openrouter.ai/openai/gpt-oss-120b
Notable capabilities:
  - Single-GPU open MoE: 117B total / 5.1B active MoE with MXFP4 weights; runs on one 80GB H100/MI300X. (https://huggingface.co/openai/gpt-oss-120b)
  - Open reasoning with full CoT: Configurable low/medium/high reasoning with full chain-of-thought access, harmony format. (https://huggingface.co/openai/gpt-oss-120b)

### gpt-oss-20b
**gpt-oss-20b** (OpenAI; current; reasoning-llm; released 2025-08-05; open weights) | ctx 131,072 | OpenAI API (docs): `gpt-oss-20b`; Azure OpenAI (Microsoft Foundry): `gpt-oss-20b`; OpenRouter: `openai/gpt-oss-20b` | Hugging Face: https://huggingface.co/openai/gpt-oss-20b — Open weights (Apache 2.0). Azure lists it as Preview. Safety-classifier variant openai/gpt-oss-safeguard-20b also on OpenRouter.
  - Hugging Face — https://huggingface.co/openai/gpt-oss-20b
  - OpenAI API (docs): `gpt-oss-20b` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-oss-20b)
  - Azure OpenAI (Microsoft Foundry): `gpt-oss-20b` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-oss-20b` — https://openrouter.ai/openai/gpt-oss-20b
Notable capabilities:
  - Laptop-class open reasoning: 21B total / 3.6B active MoE in MXFP4; runs in ~16GB memory. (https://huggingface.co/openai/gpt-oss-20b)
  - Fine-tunable on consumer hardware: Apache 2.0 weights, fine-tunable locally; function calling and structured outputs. (https://huggingface.co/openai/gpt-oss-20b)

### GPT-4o mini TTS
**GPT-4o mini TTS** (OpenAI; current; audio/speech; released 2025-03-20) | OpenAI API: `gpt-4o-mini-tts`; Azure OpenAI (Microsoft Foundry): `gpt-4o-mini-tts` — Snapshots gpt-4o-mini-tts-2025-03-20 and gpt-4o-mini-tts-2025-12-15 (default). Older tts-1 ($15/1M chars) and tts-1-hd ($30) still priced.
  - OpenAI API: `gpt-4o-mini-tts` — https://api.openai.com/v1/audio/speech (docs: https://developers.openai.com/api/docs/models/gpt-4o-mini-tts)
  - Azure OpenAI (Microsoft Foundry): `gpt-4o-mini-tts` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/models/gpt-4o-mini-tts
Notable capabilities:
  - Steerable speech: Only current OpenAI TTS model listed in the models catalog; max 2000 input tokens. (https://developers.openai.com/api/docs/models/gpt-4o-mini-tts)
  - Instruction-steerable voice: An `instructions` field controls accent, emotional range, intonation, impressions, speed, tone and whispering. (https://developers.openai.com/api/docs/guides/text-to-speech)

### text-embedding-3-large
**text-embedding-3-large** (OpenAI; current; embedding; released 2024-01-25) | $0.13 in / $? out per 1M tokens (USD) | OpenAI API: `text-embedding-3-large`; Azure OpenAI (Microsoft Foundry): `text-embedding-3-large` — Output is an embedding vector. Still OpenAI's newest embedding model as of 2026-09.
  - OpenAI API: `text-embedding-3-large` — https://api.openai.com/v1/embeddings (docs: https://developers.openai.com/api/docs/models/text-embedding-3-large)
  - Azure OpenAI (Microsoft Foundry): `text-embedding-3-large` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Multilingual embeddings: Most capable OpenAI embedding model for English and non-English tasks. (https://developers.openai.com/api/docs/models/text-embedding-3-large)
  - Shortenable (Matryoshka-style) vectors: Default 3072 dimensions; the `dimensions` API parameter truncates embeddings while keeping semantic quality. Max input 8192 tokens. (https://developers.openai.com/api/docs/guides/embeddings)

### text-embedding-3-small
**text-embedding-3-small** (OpenAI; current; embedding; released 2024-01-25) | $0.02 in / $? out per 1M tokens (USD) | OpenAI API: `text-embedding-3-small`; Azure OpenAI (Microsoft Foundry): `text-embedding-3-small` — Output is an embedding vector.
  - OpenAI API: `text-embedding-3-small` — https://api.openai.com/v1/embeddings (docs: https://developers.openai.com/api/docs/models/text-embedding-3-small)
  - Azure OpenAI (Microsoft Foundry): `text-embedding-3-small` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Cheap embeddings: Improved successor to ada-002 at $0.02 per 1M tokens. (https://developers.openai.com/api/docs/models/text-embedding-3-small)
  - Shortenable vectors: Default 1536 dimensions; can be shortened with the `dimensions` parameter. Max input 8192 tokens. (https://developers.openai.com/api/docs/guides/embeddings)

### GPT-5.6 Luna
**GPT-5.6 Luna** (OpenAI; legacy; reasoning-llm; released 2026-07-09) | ctx 1,050,000 | $0.2 in / $1.2 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.6-luna`; Azure OpenAI (Microsoft Foundry): `gpt-5.6-luna`; OpenRouter: `openai/gpt-5.6-luna` | Web app: https://chatgpt.com — Superseded by GPT-6 Luna (half the price) but still available.
  - OpenAI API: `gpt-5.6-luna` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.6-luna)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.6-luna` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.6-luna` — https://openrouter.ai/openai/gpt-5.6-luna
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Budget tier with 1M context: Fast, low-cost tier with 1.05M context and full reasoning-effort range. (https://developers.openai.com/api/docs/models/gpt-5.6-luna)
  - Replacement for gpt-5-nano/mini snapshots (found after launch): Named migration target for deprecated small GPT-5 snapshots. (https://developers.openai.com/api/docs/deprecations)

### GPT-5.6 Sol
**GPT-5.6 Sol** (OpenAI; legacy; reasoning-llm; released 2026-07-09) | ctx 1,050,000 | $4 in / $20 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.6-sol`; Azure OpenAI (Microsoft Foundry): `gpt-5.6-sol`; OpenRouter: `openai/gpt-5.6-sol` | Web app: https://chatgpt.com — GPT-5.6 flagship; the gpt-5.6 alias routes here. Superseded by GPT-6 Sol/Astra but still offered. OpenRouter lists $2/$10, lower than OpenAI list price $4/$20.
  - OpenAI API: `gpt-5.6-sol` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.6-sol)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.6-sol` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.6-sol` — https://openrouter.ai/openai/gpt-5.6-sol
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Named-tier family: GPT-5.6 introduced the Sol/Terra/Luna tier names (flagship/balanced/fast) replacing pro/mini/nano naming. (https://developers.openai.com/api/docs/changelog)
  - Max reasoning effort: Reasoning effort none/low/medium/high/xhigh/max. (https://developers.openai.com/api/docs/models/gpt-5.6-sol)
  - Fast mode long context (found after launch): Fast mode extended to long-context requests on Aug 5 2026. (https://developers.openai.com/api/docs/changelog)

### GPT-Realtime-Whisper
**GPT-Realtime-Whisper** (OpenAI; legacy; audio/speech; released 2026-05-07) | ctx 16,000 | OpenAI API: `gpt-realtime-whisper` — Still listed and not deprecated, but gpt-live-transcribe (2026-07-28, same $0.017/min) adds context and keyword hints and is what OpenAI recommends in its deprecation notices; hence marked legacy here. Language list not given in docs.
  - OpenAI API: `gpt-realtime-whisper` — wss://api.openai.com/v1/realtime (transcription sessions, v1/realtime/transcription_sessions) (docs: https://developers.openai.com/api/docs/models/gpt-realtime-whisper)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Streaming speech-to-text with tunable latency: Streams transcript deltas from live audio with a latency/accuracy trade-off setting. (https://developers.openai.com/api/docs/models/gpt-realtime-whisper)

### GPT-5.5 Pro
**GPT-5.5 Pro** (OpenAI; legacy; reasoning-llm; released 2026-04-24) | ctx 1,050,000 | $30 in / $180 out per 1M tokens (USD), no cached-input discount | OpenAI API: `gpt-5.5-pro`; OpenRouter: `openai/gpt-5.5-pro` | Web app: https://chatgpt.com — Last separately-billed "-pro" API id; for GPT-5.6/GPT-6 OpenRouter exposes pro as reasoning.mode pro. Snapshot gpt-5.5-pro-2026-04-23. Azure id not verified.
  - OpenAI API: `gpt-5.5-pro` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.5-pro)
  - OpenRouter: `openai/gpt-5.5-pro` — https://openrouter.ai/openai/gpt-5.5-pro
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Extended compute: Uses more compute per request; some requests take several minutes. Effort medium/high/xhigh. (https://developers.openai.com/api/docs/models/gpt-5.5-pro)
  - Responses/Batch only: Not available on Chat Completions. (https://developers.openai.com/api/docs/models/gpt-5.5-pro)

### GPT-5.5
**GPT-5.5** (OpenAI; legacy; reasoning-llm; released 2026-04-24) | ctx 1,050,000 | $5 in / $30 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.5`; Azure OpenAI (Microsoft Foundry): `gpt-5.5`; OpenRouter: `openai/gpt-5.5` | Web app: https://chatgpt.com — Snapshot gpt-5.5-2026-04-23. Superseded by GPT-5.6 and GPT-6; still available.
  - OpenAI API: `gpt-5.5` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.5)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.5` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.5` — https://openrouter.ai/openai/gpt-5.5
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - xhigh reasoning effort: Reasoning effort none/low/medium/high/xhigh. (https://developers.openai.com/api/docs/models/gpt-5.5)
  - 1M context: 1.05M context window with 128K output. (https://developers.openai.com/api/docs/models/gpt-5.5)

### GPT Image 2
**GPT Image 2** (OpenAI; legacy; image-gen; released 2026-04-21) | OpenAI API: `gpt-image-2`; Azure OpenAI (Microsoft Foundry): `gpt-image-2` — Snapshot gpt-image-2-2026-04-21. Superseded by GPT Image 2.5 Sunburst/Flare; still priced and not deprecated. gpt-image-1 ($10/$40 image) also still listed.
  - OpenAI API: `gpt-image-2` — https://api.openai.com/v1/images/generations (docs: https://developers.openai.com/api/docs/models/gpt-image-2)
  - Azure OpenAI (Microsoft Foundry): `gpt-image-2` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Batch image generation: Supports v1/batch in addition to generations/edits. (https://developers.openai.com/api/docs/models/gpt-image-2)
  - DALL-E replacement (found after launch): Named replacement for dall-e-2/dall-e-3 (shut down May 12 2026). (https://developers.openai.com/api/docs/deprecations)

### GPT-5.4
**GPT-5.4** (OpenAI; legacy; reasoning-llm; released 2026-03-05) | ctx 1,050,000 | $2.5 in / $15 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-5.4`; Azure OpenAI (Microsoft Foundry): `gpt-5.4`; OpenRouter: `openai/gpt-5.4` | Web app: https://chatgpt.com — Snapshot gpt-5.4-2026-03-05. Variants gpt-5.4-pro ($30/$180), gpt-5.4-mini ($0.75/$4.50), gpt-5.4-nano ($0.20/$1.25) are also on the pricing page and OpenRouter.
  - OpenAI API: `gpt-5.4` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/gpt-5.4)
  - Azure OpenAI (Microsoft Foundry): `gpt-5.4` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-5.4` — https://openrouter.ai/openai/gpt-5.4
  - Web app — https://chatgpt.com
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Tool search and computer use: Launched together with API tool search and computer-use support. (https://developers.openai.com/api/docs/changelog)
  - 1M context: 1.05M context window with 128K output; effort defaults to none. (https://developers.openai.com/api/docs/models/gpt-5.4)

### GPT-Realtime-1.5
**GPT-Realtime-1.5** (OpenAI; legacy; audio/speech; released 2026-02-23) | ctx 32,000 | OpenAI API: `gpt-realtime-1.5` — Released 2026-02-23 alongside gpt-audio-1.5 (Chat Completions). Docs still call it 'our flagship audio model for voice agents', but gpt-realtime-2 (May 2026) and gpt-realtime-2.1 (July 2026) supersede it; not deprecated as of 2026-09-29. It is the named replacement for the gpt-4o-realtime-preview models shut down 2026-05-12.
  - OpenAI API: `gpt-realtime-1.5` — wss://api.openai.com/v1/realtime (docs: https://developers.openai.com/api/docs/models/gpt-realtime-1.5)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Non-reasoning voice agent model: Speech-to-speech model for voice agents and customer support with function calling and prompt caching; cheaper text output ($16 vs $24/1M) than the reasoning gpt-realtime-2.x models. (https://developers.openai.com/api/docs/models/gpt-realtime-1.5)

### GPT-4.1
**GPT-4.1** (OpenAI; legacy; llm; released 2025-04-14) | ctx 1,047,576 | $2 in / $8 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-4.1`; Azure OpenAI (Microsoft Foundry): `gpt-4.1`; OpenRouter: `openai/gpt-4.1` — Snapshot gpt-4.1-2025-04-14. gpt-4.1-mini ($0.40/$1.60) still listed; gpt-4.1-nano deprecated, shutdown Oct 23 2026.
  - OpenAI API: `gpt-4.1` — https://api.openai.com/v1/chat/completions (docs: https://developers.openai.com/api/docs/models/gpt-4.1)
  - Azure OpenAI (Microsoft Foundry): `gpt-4.1` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-4.1` — https://openrouter.ai/openai/gpt-4.1
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - 1M-token non-reasoning model: ~1M-token context without reasoning tokens; strong instruction following and tool calling. (https://developers.openai.com/api/docs/models/gpt-4.1)
  - Fine-tunable: Supports fine-tuning, unlike the GPT-5.x models. (https://developers.openai.com/api/docs/models/gpt-4.1)

### GPT-4o
**GPT-4o** (OpenAI; legacy; multimodal; released 2024-05-13) | ctx 128,000 | $2.5 in / $10 out per 1M tokens (USD), standard tier | OpenAI API: `gpt-4o`; Azure OpenAI (Microsoft Foundry): `gpt-4o`; OpenRouter: `openai/gpt-4o` — Snapshots gpt-4o-2024-11-20, -2024-08-06, -2024-05-13 (the last deprecated, shutdown Oct 23 2026). chatgpt-4o-latest shut down Feb 17 2026. gpt-4o-mini ($0.15/$0.60) still listed.
  - OpenAI API: `gpt-4o` — https://api.openai.com/v1/chat/completions (docs: https://developers.openai.com/api/docs/models/gpt-4o)
  - Azure OpenAI (Microsoft Foundry): `gpt-4o` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/gpt-4o` — https://openrouter.ai/openai/gpt-4o
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Omni model: Natively multimodal "o" model; basis of the gpt-4o audio/realtime/transcribe/TTS variants. (https://developers.openai.com/api/docs/models/gpt-4o)
  - Fine-tunable: Supports fine-tuning via v1/fine-tuning. (https://developers.openai.com/api/docs/models/gpt-4o)

### TTS-1 / TTS-1 HD
**TTS-1 / TTS-1 HD** (OpenAI; legacy; audio/speech; released 2023-11-06) | OpenAI API: `tts-1`; OpenAI API: `tts-1-hd` — Not deprecated as of 2026-09-29 but no longer shown in the models overview; gpt-4o-mini-tts is the current, instruction-steerable replacement. Release date = OpenAI DevDay 2023 (from memory, not re-verified today).
  - OpenAI API: `tts-1` — https://api.openai.com/v1/audio/speech (docs: https://developers.openai.com/api/docs/models/tts-1)
  - OpenAI API: `tts-1-hd` — https://api.openai.com/v1/audio/speech (docs: https://developers.openai.com/api/docs/models/tts-1-hd)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Low-latency preset-voice TTS: tts-1 optimised for real-time synthesis; tts-1-hd for higher quality at twice the price. (https://developers.openai.com/api/docs/models/tts-1)

### Whisper large-v3 / large-v3-turbo (open weights)
**Whisper large-v3 / large-v3-turbo (open weights)** (OpenAI; legacy; audio/speech; released 2023-11-06; open weights) | Hugging Face: `openai/whisper-large-v3`; Hugging Face: `openai/whisper-large-v3-turbo`; Groq: `whisper-large-v3-turbo`; Deepgram (hosted): `whisper-large` | GitHub: https://github.com/openai/whisper — Status legacy: still widely deployed, but surpassed on the Open ASR Leaderboard by NVIDIA Canary/Parakeet, and OpenAI's API now points to gpt-transcribe (whisper-1 API shutdown 2027-02-26). Known to hallucinate text on silence/noise. Dates from OpenAI releases (large-v3 at DevDay 2023-11-06; turbo 2024-10-01), not re-checked today.
  - Hugging Face: `openai/whisper-large-v3` — https://huggingface.co/openai/whisper-large-v3
  - Hugging Face: `openai/whisper-large-v3-turbo` — https://huggingface.co/openai/whisper-large-v3-turbo
  - GitHub — https://github.com/openai/whisper
  - Groq: `whisper-large-v3-turbo` (docs: https://console.groq.com/docs/speech-to-text)
  - Deepgram (hosted): `whisper-large` (docs: https://developers.deepgram.com/docs/models-languages-overview)
Notable capabilities:
  - Robust multilingual ASR + translation to English: 99 languages; timestamps; zero-shot speech translation into English; the de facto open ASR baseline. (https://huggingface.co/openai/whisper-large-v3)
  - Turbo: 4-layer decoder: large-v3-turbo (Oct 2024) prunes the decoder from 32 to 4 layers (809M vs 1.55B params) for much faster decoding with minor quality loss; not trained for translation. (https://huggingface.co/openai/whisper-large-v3-turbo)

### GPT-Realtime and GPT-Realtime mini
**GPT-Realtime and GPT-Realtime mini** (OpenAI; deprecated; audio/speech; released 2025-08-28) | ctx 32,000 | OpenAI API: `gpt-realtime`; OpenAI API: `gpt-realtime-mini` — Deprecated 2026-07-20, shutdown 2027-01-20; replacements gpt-realtime-2.1 and gpt-realtime-2.1-mini. gpt-realtime-mini released 2025-10-06; its alias moved to the 2025-12-15 snapshot on 2026-01-13. Earlier gpt-4o-realtime-preview models were shut down 2026-05-12.
  - OpenAI API: `gpt-realtime` — wss://api.openai.com/v1/realtime (docs: https://developers.openai.com/api/docs/models/gpt-realtime)
  - OpenAI API: `gpt-realtime-mini` — wss://api.openai.com/v1/realtime (docs: https://developers.openai.com/api/docs/models/gpt-realtime-mini)
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - First GA OpenAI realtime speech-to-speech model: Shipped with Realtime API general availability (2025-08-28); speaks over WebRTC, WebSocket or SIP phone calls. (https://developers.openai.com/api/docs/models/gpt-realtime)

### o3
**o3** (OpenAI; deprecated; reasoning-llm; released 2025-04-16) | ctx 200,000 | $2 in / $8 out per 1M tokens (USD), standard tier | OpenAI API: `o3`; Azure OpenAI (Microsoft Foundry): `o3`; OpenRouter: `openai/o3` — Snapshot o3-2025-04-16 (and o3-pro-2025-06-10) deprecated Jun 11 2026, shutdown Dec 11 2026; replace with gpt-5.6-*. o4-mini-2025-04-16 shuts down Oct 23 2026.
  - OpenAI API: `o3` — https://api.openai.com/v1/responses (docs: https://developers.openai.com/api/docs/models/o3)
  - Azure OpenAI (Microsoft Foundry): `o3` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - OpenRouter: `openai/o3` — https://openrouter.ai/openai/o3
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Thinking with images: Reasoning model accepting image input with reasoning tokens. (https://developers.openai.com/api/docs/models/o3)
  - Successor: GPT-5 (found after launch): Docs mark o3 as succeeded by GPT-5; o-series is legacy. (https://developers.openai.com/api/docs/models/o3)

### GPT-4o Transcribe / Mini Transcribe / Transcribe Diarize
**GPT-4o Transcribe / Mini Transcribe / Transcribe Diarize** (OpenAI; deprecated; audio/speech; released 2025-03-20) | ctx 16,000 | $2.5 in / $10 out per 1M tokens (USD) for gpt-4o-transcribe and gpt-4o-transcribe-diarize (~$0.006/min); gpt-4o-mini-transcribe $1.25 / $5 (~$0.003/min) | OpenAI API: `gpt-4o-transcribe`; OpenAI API: `gpt-4o-mini-transcribe`; OpenAI API: `gpt-4o-transcribe-diarize` — Deprecated 2026-08-26, shutdown 2027-02-26 (with whisper-1); replacements gpt-transcribe (files) and gpt-live-transcribe (streaming). gpt-4o-mini-transcribe-2025-03-20 was separately deprecated 2026-07-20 in favour of the 2025-12-15 snapshot. Release date 2025-03-20 is the date of the gpt-4o-mini-tts/transcribe snapshots, not re-verified on an OpenAI launch post.
  - OpenAI API: `gpt-4o-transcribe` — https://api.openai.com/v1/audio/transcriptions (docs: https://developers.openai.com/api/docs/models/gpt-4o-transcribe)
  - OpenAI API: `gpt-4o-mini-transcribe` — https://api.openai.com/v1/audio/transcriptions (docs: https://developers.openai.com/api/docs/models/gpt-4o-mini-transcribe)
  - OpenAI API: `gpt-4o-transcribe-diarize` — https://api.openai.com/v1/audio/transcriptions
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - LLM-based transcription: Uses GPT-4o for speech-to-text with better accuracy than the original Whisper models; also usable in Realtime transcription sessions. (https://developers.openai.com/api/docs/models/gpt-4o-transcribe)

### Whisper (whisper-1 API)
**Whisper (whisper-1 API)** (OpenAI; deprecated; audio/speech; released 2023-03-01; open weights) | OpenAI API: `whisper-1` | GitHub (open weights): https://github.com/openai/whisper; Hugging Face: https://huggingface.co/openai/whisper-large-v3 — API model deprecated 2026-08-26, shutdown 2027-02-26; replacements gpt-transcribe / gpt-live-transcribe. The open-source Whisper checkpoints (MIT, first released Sept 2022) remain downloadable and widely self-hosted; the API's whisper-1 has no snapshot versions. API launch date (March 2023, with the ChatGPT API) is from memory, not re-verified today.
  - OpenAI API: `whisper-1` — https://api.openai.com/v1/audio/transcriptions (also /v1/audio/translations) (docs: https://developers.openai.com/api/docs/models/whisper-1)
  - GitHub (open weights) — https://github.com/openai/whisper
  - Hugging Face — https://huggingface.co/openai/whisper-large-v3
  - Pricing source: https://developers.openai.com/api/docs/pricing
Notable capabilities:
  - Multilingual speech recognition, translation and language ID: General-purpose ASR trained on a large diverse audio dataset; transcribes many languages and translates speech into English. (https://developers.openai.com/api/docs/models/whisper-1)

### Sora 2
**Sora 2** (OpenAI; retired; video-gen; released 2025-10-06) | OpenAI API: `sora-2`; Azure OpenAI (Microsoft Foundry): `sora-2` — OpenAI API shut down 2026-09-24 (sora-2, sora-2-pro, snapshots sora-2-2025-10-06, sora-2-2025-12-08). Azure Foundry still listed sora-2 (preview) as of 2026-09-23. Release date = first API snapshot. Resellers followed: ElevenLabs removed Sora 2 and Sora 2 Pro from its Image & Video API on 2026-09-23 ('OpenAI is discontinuing the Sora API on September 24, 2026'), and the same changelog lists ByteDance retiring Seedance 1.5 Pro on 2026-11-11.
  - OpenAI API: `sora-2` — https://api.openai.com/v1/videos (docs: https://developers.openai.com/api/docs/models/sora-2)
  - Azure OpenAI (Microsoft Foundry): `sora-2` (docs: https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
  - Pricing source: https://developers.openai.com/api/docs/models/sora-2
Notable capabilities:
  - Synchronized audio: Generates video with audio from text or image prompts. (https://developers.openai.com/api/docs/models/sora-2)
  - Shut down without replacement (found after launch): Sora 2 models and Videos API shut down Sep 24 2026 with no one-to-one replacement. (https://developers.openai.com/api/docs/deprecations)


## Physical Intelligence

### π0.7
**π0.7** (Physical Intelligence; current; robotics; released 2026-04-16) | None (internal / partner deployments; no public weights or API): https://www.pi.website/blog/pi07 — PI describes 'the first signs of compositional generalization' in its own models; not marked first:true. No weights in openpi as of 2026-09-29 (latest open PI model is π0.5). Parameter count not found. No newer PI model found through 2026-09-29.
  - None (internal / partner deployments; no public weights or API) — https://www.pi.website/blog/pi07
Notable capabilities:
  - Compositional generalization to untrained tasks: Recombines skills to do tasks never in training (e.g. operating an air fryer seen only in two fragmentary training episodes; laundry folding on a robot with no folding data). (https://www.pi.website/blog/pi07)
  - Steerable by natural-language coaching: Plain-language coaching lifted air-fryer success from ~5% to ~95% in about 30 minutes, without retraining. (https://techcrunch.com/2026/04/16/physical-intelligence-a-hot-robotics-startup-says-its-new-robot-brain-can-figure-out-tasks-it-was-never-taught/)
  - Generalist matches fine-tuned specialists: One general model performs dexterous tasks at the level of per-task fine-tuned specialists and transfers across embodiments. (https://www.pi.website/blog/pi07)

### π0.5
**π0.5** (Physical Intelligence; current; robotics; released 2025-04-22; open weights) | GitHub (openpi, JAX + PyTorch): `gs://openpi-assets/checkpoints/pi05_base`; Hugging Face (LeRobot port): `lerobot/pi05_base` — Announced 2025-04-22; weights open-sourced Sept 2025 (pi05_base, pi05_libero, pi05_droid). Still the most capable open-weights PI model. Full fine-tuning needs >70 GB VRAM.
  - GitHub (openpi, JAX + PyTorch): `gs://openpi-assets/checkpoints/pi05_base` — https://github.com/Physical-Intelligence/openpi
  - Hugging Face (LeRobot port): `lerobot/pi05_base` — https://huggingface.co/lerobot/pi05_base (docs: https://huggingface.co/docs/lerobot/en/pi05)
Notable capabilities:
  - Open-world generalization to unseen homes: Cleans kitchens and bedrooms in entirely new homes not in training; performance improved as training grew from 3 to 104 homes; co-trained on heterogeneous robot, web and verbal-instruction data. (https://www.pi.website/blog/pi05)
  - Hierarchical subtask prediction + actions in one model: Predicts a high-level text subtask, then low-level actions; trained with knowledge insulation. (https://github.com/Physical-Intelligence/openpi)
  - Newest open-weights π model (found after launch): Base plus LIBERO and DROID checkpoints released in openpi in September 2025; the latest PI model with public weights as of 2026-09 (π0.6/π0.7 are closed). (https://github.com/Physical-Intelligence/openpi)

### π0.6 / π*0.6
**π0.6 / π*0.6** (Physical Intelligence; legacy; robotics; released 2025-11-17) | None (internal to Physical Intelligence; model card and paper only): https://www.pi.website/blog/pistar06 — π0.6: ~5B-parameter VLA with a Gemma 3 4B backbone and ~860M-parameter action expert, keeps π0.5's hierarchical design (per model card, 2025-11-17, via search snippet). No weights or API. Superseded by π0.7 (2026-04). pi.website blocked automated fetches on 2026-09-29; details taken from search snippets of the blog/model card.
  - None (internal to Physical Intelligence; model card and paper only) — https://www.pi.website/blog/pistar06 (docs: https://website.pi-asset.com/pi06star/PI06_model_card.pdf)
Notable capabilities:
  - Recap - RL from real-world experience and corrections: π*0.6 improves π0.6 with Recap (RL with Experience & Corrections via Advantage-conditioned Policies): demonstrations, then human interventions, then autonomous-trial RL; over 2x throughput and roughly halved failure rates on hard tasks. (https://www.pi.website/blog/pistar06)
  - Hours-long autonomous operation: Made espresso drinks for 18 hours straight, folded 50 novel laundry items in a new home, and assembled/labeled 59 factory boxes. (https://www.pi.website/blog/pistar06)

### π0-FAST
**π0-FAST** (Physical Intelligence; legacy; robotics; released 2025-01-16; open weights) | GitHub (openpi): `gs://openpi-assets/checkpoints/pi0_fast_base`; Hugging Face (LeRobot port): `lerobot/pi0fast-base` — FAST tokenizer released and open-sourced mid-January 2025 (X post by @physical_int, 2025-01-16 approx.); π0-FAST weights open-sourced in openpi on 2025-02-04. Not marked first: no explicit 'first' claim verified.
  - GitHub (openpi): `gs://openpi-assets/checkpoints/pi0_fast_base` — https://github.com/Physical-Intelligence/openpi
  - Hugging Face (LeRobot port): `lerobot/pi0fast-base` — https://huggingface.co/lerobot/pi0fast-base (docs: https://huggingface.co/docs/lerobot/pi0fast)
Notable capabilities:
  - FAST action tokenizer (autoregressive VLA): Frequency-space Action Sequence Tokenization (DCT + BPE) compresses action chunks ~10x, letting an autoregressive VLA learn dexterous high-frequency tasks and train up to 5x faster than diffusion/flow π0. (https://huggingface.co/blog/pi0)
  - DROID generalist checkpoint (found after launch): pi0_fast_droid runs zero-shot on Franka DROID setups for many table-top instructions (openpi). (https://github.com/Physical-Intelligence/openpi)

### π0 (pi-zero)
**π0 (pi-zero)** (Physical Intelligence; legacy; robotics; released 2024-10-31; open weights) | GitHub (openpi, JAX + PyTorch): `gs://openpi-assets/checkpoints/pi0_base`; Hugging Face (LeRobot port): `lerobot/pi0_base` — Announced 2024-10-31; weights released 2025-02-04 (openpi). Not the first open VLA (OpenVLA/Octo came earlier) but became the most widely used open generalist robot policy baseline. Superseded by π0.5; still available. Fine-tuned expert checkpoints: pi0_droid, pi0_aloha_towel, pi0_aloha_tupperware, pi0_aloha_pen_uncap.
  - GitHub (openpi, JAX + PyTorch): `gs://openpi-assets/checkpoints/pi0_base` — https://github.com/Physical-Intelligence/openpi
  - Hugging Face (LeRobot port): `lerobot/pi0_base` — https://huggingface.co/lerobot/pi0_base (docs: https://huggingface.co/docs/lerobot/pi0)
Notable capabilities:
  - Flow-matching VLA for dexterous, high-frequency control: PaliGemma VLM plus an action expert that outputs continuous action chunks via flow matching (up to 50 Hz), trained on data from 8 distinct robots; folds laundry, busses tables, assembles boxes. (https://www.pi.website/blog/pi0)
  - Open weights with fine-tuning recipes (found after launch): Open-sourced 2025-02-04 in openpi with base and fine-tuned checkpoints (ALOHA towel/tupperware/pen, DROID) pre-trained on 10k+ hours of robot data; inference needs >8 GB VRAM, LoRA fine-tuning >22.5 GB. (https://github.com/Physical-Intelligence/openpi)


## Recraft

### Recraft V4.1
**Recraft V4.1** (Recraft; current; image-gen; released 2026-05-14) | Recraft API: `recraftv4_1`; OpenRouter: `recraft/recraft-v4.1` | Web app: https://www.recraft.ai — Model ids: recraftv4_1, recraftv4_1_pro, recraftv4_1_vector, recraftv4_1_pro_vector, recraftv4_1_utility(_pro)(_vector), recraftv4_1_flash; earlier recraftv4 ($0.04), recraftv4_styles, recraftv3. OpenAI-SDK compatible.
  - Recraft API: `recraftv4_1` — https://external.api.recraft.ai/v1/images/generations (docs: https://www.recraft.ai/docs/api-reference/getting-started)
  - OpenRouter: `recraft/recraft-v4.1`
  - Web app — https://www.recraft.ai
  - Pricing source: https://www.recraft.ai/docs/api-reference/pricing
Notable capabilities:
  - Native vector (SVG) generation: Dedicated Vector variants (recraftv4_1_vector, _pro_vector) output editable vector logos, typography and illustrations. (https://www.recraft.ai/blog/recraft-v4-1-more-beautiful-by-nature)
  - Utility variant for mockups: V4.1 Utility gives flat lighting, front-facing product/mockup outputs alongside the expressive main model. (https://www.recraft.ai/blog/recraft-v4-1-more-beautiful-by-nature)
  - V4.1 Flash (found after launch): Sept 2026 fast variant (~1.3 s end-to-end) at $0.007/image. (https://www.recraft.ai/docs/api-reference/getting-started)


## Resemble AI

### Resemble AI Chatterbox (Turbo / Nano / Multilingual V3)
**Resemble AI Chatterbox (Turbo / Nano / Multilingual V3)** (Resemble AI; current; audio/speech; released 2025-05-28; open weights) | Hugging Face: `ResembleAI/chatterbox`; Hugging Face: `ResembleAI/chatterbox-turbo`; pip: `chatterbox-tts`; NVIDIA NIM: `resembleai/chatterbox-multilingual-tts` — All MIT-licensed. Multilingual (23 langs) first released Sept 2025. Multilingual V3 released 2026-06-10 (Resemble post; V3 T3 weights first pushed to HF 2026-04-22): same 0.5B Llama backbone, training data up from 25.6k to 36.7k hours, 25 languages incl. 4 dialects and 6 tuned Language Pack models, PerTh watermark on by default; Resemble reports CER under 0.20% for Italian/German but ~71-75% for Korean/Vietnamese (not production-ready); NVIDIA NIM claims 2x-39x throughput. Chatterbox-Nano HF repo created 2026-04-14 (public announcement date not found). Artificial Analysis lists Chatterbox at ~1020 Elo (secondary source). Resemble's pricing page now centres on deepfake detection; hosted TTS price not verified.
  - Hugging Face: `ResembleAI/chatterbox` — https://huggingface.co/ResembleAI/chatterbox
  - Hugging Face: `ResembleAI/chatterbox-turbo` — https://huggingface.co/ResembleAI/chatterbox-turbo
  - pip: `chatterbox-tts` — https://github.com/resemble-ai/chatterbox
  - NVIDIA NIM: `resembleai/chatterbox-multilingual-tts` — https://build.nvidia.com/resembleai/chatterbox-multilingual-tts/modelcard
Notable capabilities:
  - Emotion exaggeration control: Original 0.5B Chatterbox exposes an exaggeration/intensity knob plus CFG; zero-shot cloning from ~5 s. (https://github.com/resemble-ai/chatterbox)
  - Chatterbox-Turbo: one-step decoder, paralinguistic tags (found after launch): 350M params (Dec 2025); speech-token-to-mel decoder distilled from 10 steps to 1; native [laugh], [cough], [chuckle] tags; sub-200 ms production latency. (https://huggingface.co/ResembleAI/chatterbox-turbo)
  - Built-in PerTh watermark: Every output carries Resemble's imperceptible Perth neural watermark that survives MP3 compression and edits. (https://github.com/resemble-ai/chatterbox)
  - Multilingual V3 and Nano (found after launch): Multilingual V3 (0.5B, 23 languages, better speaker similarity, fewer hallucinations) plus single-language fine-tune packs; Chatterbox-Nano (110M, English, ~3x real time on 8 CPU cores). (https://github.com/resemble-ai/chatterbox)


## Rime

### Rime Arcana v3 / v3 Turbo
**Rime Arcana v3 / v3 Turbo** (Rime; current; audio/speech; released 2026-02-04) | Rime API: `arcana`; Together AI: `Rime Arcana V3 / Arcana V3 Turbo (dedicated endpoints)` | Telnyx: https://telnyx.com/release-notes/rime-arcana-v3-voices; On-prem: https://www.rime.ai/resources/arcana-v3 — Calling the existing `arcana` model id automatically serves v3. Arcana V3 Turbo is the low-latency variant (Together AI: ~120 ms time-to-first-audio, $10 per 1M characters plus GPU-hour on dedicated endpoints). Earlier: Arcana (Apr 2025), Arcana v2. Rime's own per-character price not verified.
  - Rime API: `arcana` — https://users.rime.ai (geo endpoints users-west.rime.ai, users-east.rime.ai; HTTP and WebSocket JSON streaming) (docs: https://docs.rime.ai/)
  - Together AI: `Rime Arcana V3 / Arcana V3 Turbo (dedicated endpoints)` — https://www.together.ai/models/rime-arcana-v3-turbo
  - Telnyx — https://telnyx.com/release-notes/rime-arcana-v3-voices
  - On-prem — https://www.rime.ai/resources/arcana-v3
Notable capabilities:
  - Native code-switching across 10 languages: One voice switches mid-conversation among English, Hindi, Spanish, Arabic, French, Portuguese, German, Japanese, Hebrew and Tamil (Together AI lists 11 languages); word-level timestamps. (https://www.rime.ai/resources/arcana-v3)
  - Enterprise latency and on-prem scale: ~120 ms on-prem model latency, ~200 ms TTFB via cloud API, 100+ concurrent generations per machine; Rapidata listener tests (US) preferred it 61-64% of the time over ElevenLabs Turbo v2.5, Google Chirp and Cartesia Sonic (vendor-run). (https://www.rime.ai/resources/arcana-v3)


## Runway

### Runway Aleph 2.0
**Runway Aleph 2.0** (Runway; current; video-gen; released 2026-05-21) | Runway API: `aleph2` | Web app: https://app.runwayml.com — Launched with Edit Studio 2026-05-21; API since 2026-06-02 (2-30 s input videos). Supersedes gen4_aleph (removed from API 2026-07-30).
  - Runway API: `aleph2` — https://api.dev.runwayml.com/v1/video_to_video (docs: https://docs.dev.runwayml.com/guides/models/)
  - Web app — https://app.runwayml.com
  - Pricing source: https://docs.dev.runwayml.com/guides/pricing/
Notable capabilities:
  - In-context video editing of real footage: Edits existing clips (up to 30 s of 1080p): change angles, lighting, objects, wardrobe, background while preserving untouched motion and scene structure. (https://runway.com/news/introducing-aleph-2-and-edit-studio)
  - Edit one frame, propagate to the clip: Image-level keyframe control (up to 5 keyframes in the API) and multi-shot edits applied across scene cuts. (https://docs.dev.runwayml.com/api-details/api_changelog/)

### Runway Gen-4.5
**Runway Gen-4.5** (Runway; current; video-gen; released 2025-12-01) | Runway API: `gen4.5` | Web app: https://app.runwayml.com — Announced 2025-12-01; added to Runway API 2026-02-10 (text-to-video and image-to-video, 2-10 s). Cheaper sibling gen4_turbo (5 credits/s). gen4_aleph and gen3a_turbo removed from API 2026-07-30. Requires header X-Runway-Version: 2024-11-06.
  - Runway API: `gen4.5` — https://api.dev.runwayml.com/v1/image_to_video (docs: https://docs.dev.runwayml.com/guides/models/)
  - Web app — https://app.runwayml.com
  - Pricing source: https://docs.dev.runwayml.com/guides/pricing/
Notable capabilities:
  - #1 on Artificial Analysis text-to-video at launch: Launched as the top model on the Artificial Analysis Text-to-Video leaderboard (1,247 Elo), with better physics (liquids, momentum, collisions). (https://runway.com/research/introducing-runway-gen-4.5)
  - HDR and professional output formats (found after launch): API can output ProRes, PNG/EXR sequences, 10-bit SDR and HDR10/HLG/ACEScg masters (Gen-4.5 only for HDR). (https://docs.dev.runwayml.com/guides/models/)


## Sesame

### Sesame CSM-1B (Conversational Speech Model)
**Sesame CSM-1B (Conversational Speech Model)** (Sesame; current; audio/speech; released 2025-03-13; open weights) | Hugging Face: `sesame/csm-1b`; Transformers: `sesame/csm-1b` | Sesame app (Maya, Miles, Simone, Charlie — hosted larger models): https://www.sesame.com/ — Open base generation model only (no fine-tuned voices, English-centric, cannot generate text itself); the Maya/Miles demo voices use Sesame's larger in-house models. Native in Transformers since v4.52.1. Sesame raised a $250M Series B (Oct 2025, Sequoia/Spark) and launched a public-preview iOS app with four agents (Maya, Miles, Simone, Charlie) in 39 countries on 2026-05-28; smart glasses targeted for 2027. No newer open Sesame model found as of 2026-09-29.
  - Hugging Face: `sesame/csm-1b` — https://huggingface.co/sesame/csm-1b
  - Transformers: `sesame/csm-1b` (docs: https://huggingface.co/docs/transformers/model_doc/csm)
  - Sesame app (Maya, Miles, Simone, Charlie — hosted larger models) — https://www.sesame.com/
Notable capabilities:
  - Context-conditioned conversational TTS: Llama backbone + audio decoder emitting Mimi audio codes; generates speech conditioned on prior conversation audio/text so prosody fits the dialogue; voice prompting via context segments. (https://huggingface.co/sesame/csm-1b)


## Skild AI

### Skild S1 (Skild Brain)
**Skild S1 (Skild Brain)** (Skild AI; current; robotics; released 2026-08-25) | Skild AI (commercial partners; early-access sign-up): https://www.skild.ai/blogs/s1 — Announced on X 2026-08-25 (https://x.com/SkildAI/status/2092300842900865389); press 2026-08-31; NVIDIA blog 2026-09-10 (https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/) cites a $100M revenue run rate 10 months after first commercial deployment, 60+ deployment partnerships and Blackwell assembly work with Foxconn. Skild raised a $1.4B Series C at >$14B (2026-01-14, led by SoftBank). No public API, pricing or weights; company says S1 is "already at work with our commercial partners" and plans wider real-world rollout by 2027. Results are company-reported.
  - Skild AI (commercial partners; early-access sign-up) — https://www.skild.ai/blogs/s1
Notable capabilities:
  - [FIRST] In-context learning from one video, long-horizon: Learns tasks never seen in pretraining (potting a plant, cooking pancakes, pour-over coffee, kit assembly) from a single video prompt with no fine-tuning, for tasks up to ~10 minutes long; Skild calls this the first robotics foundation model to show in-context learning on such long unseen tasks. (https://www.skild.ai/blogs/s1)
  - Video prompting beats language prompting: 66% success on unseen tasks vs 9% for an equivalently trained language-prompted policy (~7x); 96% on seen tasks; one demo video worth ~380 post-training episodes; 11 minutes from demonstration to autonomous execution in the plant-potting example. (https://www.skild.ai/blogs/s1)
  - Omni-bodied brain: Skild Brain is pitched as one model controlling quadrupeds, humanoids, arms and mobile manipulators without prior knowledge of the body; S1 trains on teleop, human video, simulation and data-capture gloves. (https://www.therobotreport.com/skild-ai-unveils-s1-flagship-robot-foundation-model/)


## Soniox

### Soniox TTS v2
**Soniox TTS v2** (Soniox; current; audio/speech; released 2026-08-10) | $4 in / $21.5 out USD per 1M tokens (text in / audio out); ≈ $0.70 per hour of generated speech (1 hour ≈ 30,000 audio tokens) | Soniox API (real-time streaming, WebSocket): `tts-rt-v2` — Soniox launched TTS on 2026-04-23 (tts-rt-v1); TTS v2 (tts-rt-v2, replacing v1) was reported by audioXpress on 2026-08-10. Streaming only; regions US, EU, Japan. The v2 date is from secondary press, not a Soniox post.
  - Soniox API (real-time streaming, WebSocket): `tts-rt-v2` (docs: https://soniox.com/text-to-speech)
  - Pricing source: https://soniox.com/pricing
Notable capabilities:
  - 60+ languages in one model, mid-sentence switching: Single multilingual model with mixed-language text and mid-sentence language switching; Soniox claims 'hallucination-free' output (no invented or dropped words) and accurate reading of emails, phone numbers and IDs. (https://soniox.com/blog/soniox-text-to-speech)
  - Audio tags and 20-second voice cloning (v2): TTS v2 adds expressive audio tags (whispering, laughter, hesitation, excitement), voice cloning from ~20 s of reference audio, and character-level timestamps. (https://audioxpress.com/news/soniox-tts-v2-adds-expressive-control-and-voice-cloning-to-its-multilingual-voice-ai-platform)

### Soniox v5 (Async and Real-Time STT)
**Soniox v5 (Async and Real-Time STT)** (Soniox; current; audio/speech; released 2026-06-11) | $1.5 in / $3.5 out USD per 1M tokens for async (audio in $1.50, text in/out $3.50; ~$0.10 per audio hour). Real-time: $2.00 audio in, $4.00 text in/out (~$0.12/hour). 1 hour of audio ≈ 30,000 input tokens. | Soniox API (async / file): `stt-async-v5`; Soniox API (real-time streaming): `stt-rt-v5` | Web app: https://soniox.com — stt-async-v5 released 2026-06-11, stt-rt-v5 on 2026-06-16. The v4 ids (stt-async-v4 from 2026-01-29, stt-rt-v4 from 2026-02-05) were retired 2026-06-30 and are now aliases routing to v5. Launch posts give no WER numbers; Soniox publishes its own comparisons at soniox.com/benchmarks (vendor-run). Sibling TTS: soniox-tts-v2.
  - Soniox API (async / file): `stt-async-v5` (docs: https://soniox.com/docs/stt/models)
  - Soniox API (real-time streaming): `stt-rt-v5` (docs: https://soniox.com/docs/stt/models)
  - Web app — https://soniox.com
  - Pricing source: https://soniox.com/pricing
Notable capabilities:
  - One multilingual model for 60+ languages with speaker separation: Soniox claims native-speaker accuracy across 60+ languages in a single model, re-engineered speaker diarization, spoken-language ID, context injection and precise alphanumerics (IDs, emails, codes). (https://soniox.com/blog/soniox-v5-async)
  - Real-time translation and semantic endpointing: stt-rt-v5 transcribes and translates live across ~3,600 language pairs, with a tunable `endpoint_sensitivity` semantic endpointing parameter for voice agents. (https://soniox.com/blog/soniox-v5-real-time)


## Speechify (SpeechifyAI)

### Speechify Simba 3.2
**Speechify Simba 3.2** (Speechify (SpeechifyAI); current; audio/speech; released 2026-07-07) | Web: https://speechify.ai/models — Exact API model id string not verified (docs page 'SpeechifyAI Build TTS Models: Simba 3.2, 3.0, Multilingual, and English'). AA measured ~30.2 chars/s generation speed (the-decoder, Jul 2026). Quotes: Luke Oliff, Tyler Weitzman in the press release.
  - SpeechifyAI API — https://api.speechify.ai/v1/audio/speech (also /v1/audio/stream) (docs: https://docs.speechify.ai/build/guides/concepts/models)
  - Web — https://speechify.ai/models
  - Pricing source: https://www.prweb.com/releases/speechifys-simba-3-2-ranks-1-on-independent-artificial-analysis-tts-leaderboard-worlds-best-real-time-voice-model-above-elevenlabs-openai-google-deepmind--others-302819731.html
Notable capabilities:
  - Briefly #1 on Artificial Analysis Speech Arena at a low price: Press release 2026-07-07 claimed #1 on the AA TTS leaderboard; a week later Qwen-Audio-3.0-TTS-Plus overtook it (1,236 vs 1,234 Elo). On 2026-09-29 it was #7 (Elo 1239). Speechify called it the cheapest model in the top ten ($10/$6 per 1M chars). (https://artificialanalysis.ai/text-to-speech/leaderboard)
  - Streaming-native, low TTFB: Streaming-native Simba 3 model; <100 ms first byte claimed; emotional control, SSML prosody, instant voice cloning; 30+ locales with mixed-language input. Recommended model for English integrations. (https://speechify.ai/blog/simba-3-2-streaming-model)


## Speechmatics

### Speechmatics Linden 1 (Agent STT)
**Speechmatics Linden 1 (Agent STT)** (Speechmatics; current; audio/speech; released 2026-09-17) | Speechmatics Agent STT API: `linden-1` | Pipecat: https://www.speechmatics.com/voice-agents; LiveKit: https://docs.livekit.io/agents/models/stt/speechmatics/ — Targets high-consequence errors in calls (a changed digit, a missed 'not', a one-word confirmation). Benchmark figures are vendor-reported from Pipecat's public benchmark. Sibling batch model: speechmatics-melia-1.
  - Speechmatics Agent STT API: `linden-1` — /v2/agent (regions eu1 / us1 .asr.api.speechmatics.com) (docs: https://docs.speechmatics.com/speech-to-text/models)
  - Pipecat — https://www.speechmatics.com/voice-agents
  - LiveKit — https://docs.livekit.io/agents/models/stt/speechmatics/
  - Pricing source: https://www.globenewswire.com/news-release/2026/09/17/3364138/0/en/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents.html
Notable capabilities:
  - STT output shaped for LLM voice agents: Returns speaker-attributed segments with turn messages instead of a running word stream; finalizes segments in under 350 ms; 55+ languages; custom vocabulary up to 1,000 terms; live diarization and speaker ID. (https://docs.speechmatics.com/speech-to-text/models)
  - Low semantic error on Pipecat benchmark: 1.05% pooled semantic error rate and 369 ms median finalization on the Pipecat STT benchmark (23 streaming models), on the speed/accuracy Pareto frontier, per Speechmatics. (https://www.globenewswire.com/news-release/2026/09/17/3364138/0/en/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents.html)

### Speechmatics Melia 1 (multilingual STT)
**Speechmatics Melia 1 (multilingual STT)** (Speechmatics; preview; audio/speech; released 2026-06-17) | Speechmatics Batch API: `melia-1` — Launched 2026-06-17 as a production preview (docs: early access), batch only; runs alongside the Standard and Enhanced models. Benchmarks are vendor-reported.
  - Speechmatics Batch API: `melia-1` — batch jobs with "model": "melia-1" and "language": "multi" (EU1, US1) (docs: https://docs.speechmatics.com/speech-to-text/models)
  - Pricing source: https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model
Notable capabilities:
  - Code-switching across 55+ languages without language selection: Transcribes audio that switches languages mid-conversation with no language pre-selection; Speechmatics reports it beats Deepgram and Microsoft on 91% and AssemblyAI on 77% of FLEURS languages, and 5% lower WER than its Standard model on noisy monolingual audio. (https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model)


## Stability AI

### Stable Audio 3.0
**Stable Audio 3.0** (Stability AI; current; music; released 2026-05-20; open weights) | Hugging Face (Medium): https://huggingface.co/stabilityai/stable-audio-3-medium; Hugging Face (Small music): https://huggingface.co/stabilityai/stable-audio-3-small-music; Hugging Face (Small SFX): https://huggingface.co/stabilityai/stable-audio-3-small-sfx; Web app: https://stableaudio.com — Family of 4: Small SFX, Small, Medium (open weights, HF) and Large (API via Stability and fal.ai, or enterprise self-hosting). Exact API model id/endpoint for Large not verified (Stability pricing/docs pages are JS-rendered). Price per The Rundown tool review (says it checked the official pricing page 2026-08-31, secondary): 26 API credits = $0.26 per successful Large generation (1 credit = $0.01).
  - Stability AI API (Large) (docs: https://platform.stability.ai)
  - Hugging Face (Medium) — https://huggingface.co/stabilityai/stable-audio-3-medium
  - Hugging Face (Small music) — https://huggingface.co/stabilityai/stable-audio-3-small-music
  - Hugging Face (Small SFX) — https://huggingface.co/stabilityai/stable-audio-3-small-sfx
  - Web app — https://stableaudio.com
Notable capabilities:
  - Tracks over 6 minutes: Medium generates music up to 6:20; Large aimed at high-volume, low-latency platform use. (https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models)
  - Fully licensed training data: Model family trained on fully licensed data; users own outputs under the Community License. (https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models)
  - On-device small models: Small (459M) music and Small SFX models designed to run on phones and consumer laptops. (https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models)

### Stable Diffusion 3.5 Large
**Stable Diffusion 3.5 Large** (Stability AI; current; image-gen; released 2024-10-22; open weights) | Stability AI API: `sd3.5-large` | Hugging Face: https://huggingface.co/stabilityai/stable-diffusion-3.5-large; Hugging Face (Large Turbo): https://huggingface.co/stabilityai/stable-diffusion-3.5-large-turbo; Hugging Face (Medium): https://huggingface.co/stabilityai/stable-diffusion-3.5-medium — Still Stability's latest image model family (no official SD4 as of 2026-09; SD4 'news' articles are unverified). API model values for /generate/sd3: sd3.5-large, sd3.5-large-turbo, sd3.5-medium (from third-party docs; official API ref is JS-rendered, not verified). Pricing (credits) not verified.
  - Stability AI API: `sd3.5-large` — https://api.stability.ai/v2beta/stable-image/generate/sd3 (docs: https://platform.stability.ai/docs/api-reference)
  - Hugging Face — https://huggingface.co/stabilityai/stable-diffusion-3.5-large
  - Hugging Face (Large Turbo) — https://huggingface.co/stabilityai/stable-diffusion-3.5-large-turbo
  - Hugging Face (Medium) — https://huggingface.co/stabilityai/stable-diffusion-3.5-medium
Notable capabilities:
  - Open MMDiT weights, free for small businesses: 8B Multimodal Diffusion Transformer with open weights under a Community License free for commercial use under $1M annual revenue. (https://huggingface.co/stabilityai/stable-diffusion-3.5-large)
  - Broad hardware optimization (found after launch): Official TensorRT/FP8 (NVIDIA, ~2x faster, 40% less memory), ONNX AMD GPU and AMD NPU builds released later. (https://stability.ai/news-updates)


## Stanford / UC Berkeley / Toyota Research Institute

### OpenVLA (7B) and OpenVLA-OFT
**OpenVLA (7B) and OpenVLA-OFT** (Stanford / UC Berkeley / Toyota Research Institute; legacy; robotics; released 2024-06-13; open weights) | Hugging Face: `openvla/openvla-7b`; Hugging Face (OFT fine-tunes): `moojink/openvla-7b-oft-finetuned-libero-spatial` | GitHub: https://github.com/openvla/openvla — The most-downloaded open VLA checkpoint on HF (500k+ downloads at check time); widely used as a research baseline. Superseded in capability by pi0-family and newer open VLAs but still a standard reference. Release day: arXiv 2406.09246 v1 dated 2024-06-13 (HF repo created 2024-06-10).
  - Hugging Face: `openvla/openvla-7b` — https://huggingface.co/openvla/openvla-7b
  - Hugging Face (OFT fine-tunes): `moojink/openvla-7b-oft-finetuned-libero-spatial` — https://huggingface.co/moojink/openvla-7b-oft-finetuned-libero-spatial
  - GitHub — https://github.com/openvla/openvla (docs: https://openvla.github.io/)
Notable capabilities:
  - Open 7B generalist VLA beating a 55B closed model: Llama 2 7B backbone with fused DINOv2 + SigLIP vision, trained on ~970k Open X-Embodiment episodes; outperformed RT-2-X (55B) by 16.5% absolute success over 29 tasks with 7x fewer parameters, and fine-tunes with LoRA on consumer GPUs. (https://arxiv.org/abs/2406.09246)
  - OFT fine-tuning recipe (Feb 2025) (found after launch): OpenVLA-OFT (parallel decoding, action chunking, continuous actions, L1 loss) raised LIBERO average success from 76.5% to 97.1% and action throughput 26x; on bimanual ALOHA it beat pi0 and RDT-1B by up to 15% absolute. (https://arxiv.org/abs/2502.19645)


## StepFun

### StepAudio 3 ASR Max / StepAudio 3 TTS
**StepAudio 3 ASR Max / StepAudio 3 TTS** (StepFun; current; audio/speech; released 2026-09-15) | StepFun API: `stepaudio-3-asr-max`; StepFun API: `stepaudio-3-tts`; StepFun API (preview): `stepaudio-3-gen-preview` — Family file for the non-realtime StepAudio 3 models. Languages: zh, en, ja, ko, fr, es (non-zh/en in preview). stepaudio-3-gen-preview (speech+SFX+ambience+BGM) and stepaudio-3-music-preview are free during preview. Previous gen: stepaudio-2.5-asr ($0.022/h), stepaudio-2.5-asr-stream ($0.18/h), stepaudio-2.5-tts ($0.85/10k chars). Exact per-model release date assumed = family launch 2026-09-15.
  - StepFun API: `stepaudio-3-asr-max` (docs: https://platform.stepfun.ai/docs/en/guides/models/audio)
  - StepFun API: `stepaudio-3-tts` (docs: https://platform.stepfun.ai/docs/en/guides/models/audio)
  - StepFun API (preview): `stepaudio-3-gen-preview` (docs: https://platform.stepfun.ai/docs/en/guides/models/audio)
  - Pricing source: https://platform.stepfun.ai/docs/en/pricing/details
Notable capabilities:
  - #1 non-streaming ASR on AA-WER: Artificial Analysis ranked StepAudio 3 ASR #1 on its AA-WER Index for non-streaming speech-to-text with 1.7% WER (StepAudio 2.5 ASR: 4.7%). (https://x.com/ArtificialAnlys/status/2102485740248842710)
  - Context-aware streaming TTS: Natural, context-aware speech with low-latency streaming, natural-language control and voice cloning; 1,000-char input limit; wav/mp3/flac/opus/pcm. (https://platform.stepfun.ai/docs/en/guides/models/audio)

### StepFun Step-Audio-EditX
**StepFun Step-Audio-EditX** (StepFun; current; audio/speech; released 2025-11-06; open weights) | Hugging Face: `stepfun-ai/Step-Audio-EditX`; Hugging Face (4-bit): `stepfun-ai/Step-Audio-EditX-AWQ-4bit` | GitHub: https://github.com/stepfun-ai/Step-Audio-EditX — Official changelog lists a new model release on 2026-01-29 (overall ~4% improvement; new paralinguistic tags such as exhale, inhale, chuckle, clears throat, giggle; SFT/DPO/GRPO training code released); HF weights updated 2026-01-23/24, README edits to 2026-02-14. No March 2026 release appears in the official GitHub/HF changelog, so Artificial Analysis's 'Step Audio EditX (Mar 2026)' label (#3 open weights, ~1095 Elo, Sept 2026) probably refers to the Jan 2026 weights or a hosted snapshot (unverified).
  - Hugging Face: `stepfun-ai/Step-Audio-EditX` — https://huggingface.co/stepfun-ai/Step-Audio-EditX
  - Hugging Face (4-bit): `stepfun-ai/Step-Audio-EditX-AWQ-4bit` — https://huggingface.co/stepfun-ai/Step-Audio-EditX-AWQ-4bit
  - GitHub — https://github.com/stepfun-ai/Step-Audio-EditX
Notable capabilities:
  - Iterative LLM-based audio editing: 3B RL-trained audio LLM that edits emotion, speaking style and paralinguistics of existing speech step by step, plus zero-shot TTS cloning (Mandarin, English, Sichuanese, Cantonese; Japanese/Korean added 2025-11-28). (https://github.com/stepfun-ai/Step-Audio-EditX)

### StepAudio 3 Realtime
**StepAudio 3 Realtime** (StepFun; preview; audio/speech; released 2026-09-15) | $0 in / $0 out free during limited-time preview (successor stepaudio-2.5-realtime: $1.50 in / $0.30 cached / $10.00 out per 1M tokens) | StepFun API (Realtime WebSocket): `stepaudio-3-realtime-preview`; StepFun API (Chat Completions): `stepaudio-3-chat-preview` — Chinese and English. Preview ids will be retired for a paid GA version when the trial ends. Predecessor stepaudio-2.5-realtime (2026-05-26; persona role-play, project page https://stepaudiollm.github.io/step-audio-2.5-realtime/ with self-reported 86.36 general dialogue / 79.80 spoken QA / 82.18 paralinguistics, claimed to beat GPT-Realtime-1.5 on StepFun's evals). Technical report arXiv 2609.14005 (56.0% task success on tau-Voice).
  - StepFun API (Realtime WebSocket): `stepaudio-3-realtime-preview` (docs: https://platform.stepfun.ai/docs/en/guides/models/stepaudio-3-realtime)
  - StepFun API (Chat Completions): `stepaudio-3-chat-preview` (docs: https://platform.stepfun.ai/docs/en/guides/models/stepaudio-3-realtime)
  - Pricing source: https://platform.stepfun.ai/docs/en/pricing/details
Notable capabilities:
  - Think-while-speaking full duplex: Runs private chain-of-thought in parallel with spoken output; distinguishes real interruptions from backchannels; asynchronous tool execution (web search, knowledge retrieval). (https://arxiv.org/abs/2609.14005)
  - #1 on Artificial Analysis conversational dynamics: 98.9 on Artificial Analysis Full-Duplex Bench (Conversational Dynamics) and 99.7% Speech Reasoning at launch, ahead of Qwen Audio 3.0 Realtime Plus and GPT-Live-1 per StepFun. (https://x.com/StepFun_ai/status/2099916376274313630)


## Suno

### Suno v6 (v6, v6-wild, v6-mini)
**Suno v6 (v6, v6-wild, v6-mini)** (Suno; current; music; released 2026-09-09) | Web app: `v6`; Web app (Pro/Premier): `v6-wild`; Web app (all users, incl. free): `v6-mini` — Launched 2026-09-09; Suno retired all earlier models (v4 to v5.5) as v6 rolled out. No official public API: in July 2026 Suno's CPO Jack Brody announced it was only 'exploring' a developer API/partner program (intake form, no timeline); third-party 'Suno APIs' are unofficial. Monthly-billing prices ($10/$30) are derived from the pricing page's annual price ($8/$24 per month) and its stated 20% annual discount. Max song length for v6 not stated on official pages checked. Sony Music and UMG sued again on 2026-09-18 over v6.
  - Web app: `v6` — https://suno.com
  - Web app (Pro/Premier): `v6-wild`
  - Web app (all users, incl. free): `v6-mini`
  - Pricing source: https://suno.com/pricing
Notable capabilities:
  - Trained only on licensed music: First Suno generation developed with rightsholders; trained from scratch on music licensed from Warner Music Group, BMG and Believe (not on data used for earlier Suno versions), with revenue sharing to partners. (https://suno.com/blog/introducing-v6)
  - Three-variant lineup: v6 (reliable, steerable flagship), v6-wild (experimental, genre-blending, pushes away from the prompt), v6-mini (fast, high-volume, available to everyone). (https://suno.com/blog/introducing-v6)
  - Natural-language section and lyric editing: Edit parts of a song or change individual lyric lines by prompt without regenerating the whole track. (https://suno.com/blog/introducing-v6)
  - Multimodal references and mashups: Text, audio, image and video references as a starting point; combine elements of several songs into a mashup; sample/isolate instruments and build beats. (https://suno.com/blog/introducing-v6)
  - Upload screening and download limits: Uploaded audio and lyrics are screened for unauthorized use; downloads are capped per plan (none on Free, 20/month Pro, 60/month Premier). (https://suno.com/pricing)

### Suno v5.5
**Suno v5.5** (Suno; retired; music; released 2026-03-26) | Web app: https://suno.com — No official public API (web/mobile app only; third-party 'Suno APIs' are unofficial). Retired on 2026-09-09 when Suno moved entirely to the v6 family (see suno-v6); Voices and Custom Models features carried over to v6 plans.
  - Web app — https://suno.com
Notable capabilities:
  - Voices (sing with your own voice): Record/upload your voice (with verification and privacy controls) and have Suno sing songs in it; Pro/Premier. (https://about.suno.com/blog/v5-5)
  - Custom Models: Fine-tune a personal v5.5 on your own catalog (min. 6 tracks, up to 3 models per user); Pro/Premier. (https://about.suno.com/blog/v5-5)
  - My Taste personalization: Learns preferred genres/moods and applies them via the Magic Wand; all users. (https://about.suno.com/blog/v5-5)


## Tencent AI Lab

### SongGeneration 2 (LeVo 2)
**SongGeneration 2 (LeVo 2)** (Tencent AI Lab; current; music; released 2026-03-01; open weights) | Hugging Face (v2-large checkpoint, uploader account): https://huggingface.co/lglg666/SongGeneration-v2-large; Hugging Face (official org repo; returned 401 on 2026-09-29): https://huggingface.co/tencent/SongGeneration — Released 2026-03-01 (per vLLM-Omni model request citing the official repo). Reported lyric accuracy PER 8.55% vs Suno v5 12.4% and Mureka v8 9.96% (secondary source gaga.art, not verified). As of 2026-09-29 the official GitHub repo github.com/tencent-ailab/SongGeneration returns 404 and the tencent/SongGeneration HF repo returns 401 (apparently removed/made private; community forks and reuploads exist, e.g. Pinokio notes); lglg666/SongGeneration-v2-large (created 2026-02-15, license 'unknown') is still public. Treat availability and license as unverified. Demo: https://levo-demo.github.io/levo_v2_demo/
  - Hugging Face (v2-large checkpoint, uploader account) — https://huggingface.co/lglg666/SongGeneration-v2-large
  - Hugging Face (official org repo; returned 401 on 2026-09-29) — https://huggingface.co/tencent/SongGeneration
Notable capabilities:
  - Hybrid LLM-diffusion full songs up to 4:30: 4B-parameter model generating complete songs up to 4 min 30 s with vocals + accompaniment, instrumental-only, a cappella or dual-track (separated) output; multilingual lyrics (Chinese, English, Spanish, Japanese and more). (https://github.com/vllm-project/vllm-omni/issues/3390)
  - Hierarchical semantic planning + track-specific refinement (found after launch): LeVo 2 paper: semantic planning precedes per-track refinement to keep vocal-instrument coordination while improving acoustics; progressive post-training with automatic quality tiers. (https://arxiv.org/abs/2606.30642)


## Tesla

### Tesla Optimus AI (end-to-end robot neural network)
**Tesla Optimus AI (end-to-end robot neural network)** (Tesla; preview; robotics; released 2024) | Not available (internal only): https://www.tesla.com/AI — Not a product you can call: Tesla has published no model name, architecture, paper, API or weights for the Optimus neural net; this file tracks the robot AI stack. Hardware status (as of 2026-09-29): Optimus V3 / Gen 3 has NOT been unveiled. Tesla missed its Q1 2026 and "mid-2026" reveal targets; Musk said on 2026-04-22 it "will be unveiled closer to production start" and that Tesla is holding back demos because competitors copy them frame by frame. Tesla's Q1 2026 update says Fremont (former Model S/X line) is being fitted for a 1M-robot/yr first-generation line, with a Giga Texas line targeting 10M/yr long term from 2027. Rumoured V3 specs (22-DoF hands, ~$20-30K price, public sale end-2027) come from secondary sources and are unverified. Sources: Tesla Q1 2026 update and earnings call via https://en.wikipedia.org/wiki/Optimus_(robot) ; https://electrek.co/2026/04/22/tesla-optimus-production-fremont-model-sx-line/ ; https://driveteslacanada.ca/news/tesla-delaying-optimus-v3-reveal-fears-copycats/
  - Not available (internal only) — https://www.tesla.com/AI
Notable capabilities:
  - Camera-only end-to-end policy on the FSD computer: Tesla-published Optimus demos (e.g. battery-cell sorting) are described as a single end-to-end neural network running on the robot's onboard FSD computer from camera (and touch) input; Tesla shares the vision/AI stack with FSD. (https://en.wikipedia.org/wiki/Optimus_(robot))
  - Offline autonomy on AI5, Grok for conversation (found after launch): On the Q1 2026 call (2026-04-22) Musk said the AI5 chip should give Optimus enough local intelligence to keep working without connectivity, while Grok-level conversation needs WiFi/cellular. (https://en.wikipedia.org/wiki/Optimus_(robot))


## Tsinghua University (TSAIL, thu-ml)

### RDT2 (and RDT-1B)
**RDT2 (and RDT-1B)** (Tsinghua University (TSAIL, thu-ml); current; robotics; released 2025-09; open weights) | Hugging Face: `robotics-diffusion-transformer/RDT2-VQ`; Hugging Face (RDT-1B, MIT): `robotics-diffusion-transformer/rdt-1b` | GitHub: https://github.com/thu-ml/RDT2 — 'First' claim is the authors' own hedged wording. HF RDT2-VQ repo created 2025-09-22.
  - Hugging Face: `robotics-diffusion-transformer/RDT2-VQ` — https://huggingface.co/robotics-diffusion-transformer/RDT2-VQ
  - Hugging Face (RDT-1B, MIT): `robotics-diffusion-transformer/rdt-1b` — https://huggingface.co/robotics-diffusion-transformer/rdt-1b
  - GitHub — https://github.com/thu-ml/RDT2 (docs: https://rdt-robotics.github.io/rdt2/)
Notable capabilities:
  - [FIRST] Zero-shot deployment on unseen embodiments: RDT2 (8B, Qwen2.5-VL-7B based, residual-VQ action tokens; RDT2-FM flow-matching variant) trained on 10k+ h of UMI-gripper human manipulation from 100+ scenes; authors call it possibly the first foundation model to deploy zero-shot on unseen embodiments (UR5e, Franka FR3) for simple open-vocabulary tasks. (https://huggingface.co/robotics-diffusion-transformer/RDT2-VQ)
  - Large diffusion foundation model for bimanual manipulation (RDT-1B): RDT-1B (Oct 2024, 1.2B) was billed as the largest diffusion-based foundation model for bimanual manipulation, pretrained on 46 datasets (1M+ episodes) and fine-tuned on a 6K+ episode ALOHA dataset. (https://arxiv.org/abs/2410.07864)


## UC Berkeley (RAIL) / Stanford / CMU / Google DeepMind

### Octo (Octo-Small / Octo-Base 1.5)
**Octo (Octo-Small / Octo-Base 1.5)** (UC Berkeley (RAIL) / Stanford / CMU / Google DeepMind; legacy; robotics; released 2024-05-20; open weights) | Hugging Face: `rail-berkeley/octo-base-1.5` | GitHub: https://github.com/octo-models/octo — Early (2024) fully open generalist robot policy; now mostly a baseline. Parameter sizes from the project page.
  - Hugging Face: `rail-berkeley/octo-base-1.5` — https://huggingface.co/rail-berkeley/octo-base-1.5
  - GitHub — https://github.com/octo-models/octo (docs: https://octo-models.github.io/)
Notable capabilities:
  - Open generalist policy on Open X-Embodiment: Transformer diffusion policy (27M Small / 93M Base) trained on 800k trajectories from Open X-Embodiment; instructed by language or goal images; evaluated on 9 robot platforms; fine-tunes to new sensors and action spaces in hours on consumer GPUs. (https://arxiv.org/abs/2405.12213)


## Unitree Robotics

### UnifoLM-WLA-1.0
**UnifoLM-WLA-1.0** (Unitree Robotics; current; robotics; released 2026-09-10; open weights) | Hugging Face: `unitreerobotics/UnifoLM-WLA-1.0-Base`; Hugging Face (embodied reasoner backbones): `unitreerobotics/UnifoLM-ER-Flow` | GitHub: https://github.com/unitreerobotics/unifolm-wla — Staged release: announcement + demo video 2026-09-10; UnifoLM-ER-1 / ER-Flow weights 2026-09-11; model modules and training code 2026-09-20; WLA-1.0-Base weights and fine-tuning code 2026-09-28 (GitHub news). HF repo lists Apache-2.0 but the model card was empty at check time. Predecessors: UnifoLM-VLA-0 (see unifolm-vla-0) and UnifoLM-WMA-0 world-model-action (Sept 2025). Benchmark claims ("leading results across multiple embodied reasoning benchmarks") are self-reported.
  - Hugging Face: `unitreerobotics/UnifoLM-WLA-1.0-Base` — https://huggingface.co/unitreerobotics/UnifoLM-WLA-1.0-Base
  - Hugging Face (embodied reasoner backbones): `unitreerobotics/UnifoLM-ER-Flow` — https://huggingface.co/unitreerobotics/UnifoLM-ER-1
  - GitHub — https://github.com/unitreerobotics/unifolm-wla (docs: https://unigen-x.github.io/unifolm-wla.github.io/)
Notable capabilities:
  - One weight set for tabletop and whole-body humanoid manipulation: 6B-parameter model coordinating 64 tasks across tabletop and whole-body manipulation on Unitree G1, with two-finger grippers and several five-finger dexterous hands. (https://github.com/unitreerobotics/unifolm-wla)
  - Embodied reasoner + MMDiT action expert: Built on UnifoLM-ER (4B embodied reasoner based on Qwen3-VL-4B; 5M+ embodied reasoning samples) with an MMDiT action expert; ~2,500 h of real-robot data. (https://unigen-x.github.io/unifolm-wla.github.io/)

### UnifoLM-VLA-0 (and UnifoLM-WMA-0)
**UnifoLM-VLA-0 (and UnifoLM-WMA-0)** (Unitree Robotics; legacy; robotics; released 2026-01; open weights) | Hugging Face: `unitreerobotics/UnifoLM-VLA-Base`; Hugging Face (world-model-action): `unitreerobotics/UnifoLM-WMA-0-Base` | GitHub: https://github.com/unitreerobotics/unifolm-vla — HF repos for UnifoLM-VLA-Base created 2026-01-28 (exact announcement day not verified). VLA-0 license CC BY-NC-SA 4.0 (non-commercial); WMA-0 Apache-2.0. Superseded by UnifoLM-WLA-1.0 (2026-09). Unitree also publishes ~200 G1 teleoperation datasets under huggingface.co/unitreerobotics.
  - Hugging Face: `unitreerobotics/UnifoLM-VLA-Base` — https://huggingface.co/collections/unitreerobotics/unifolm-vla-0
  - Hugging Face (world-model-action): `unitreerobotics/UnifoLM-WMA-0-Base` — https://huggingface.co/unitreerobotics/UnifoLM-WMA-0-Base
  - GitHub — https://github.com/unitreerobotics/unifolm-vla
Notable capabilities:
  - Open VLA for general-purpose humanoid manipulation: Continued pretraining of a VLM (UnifoLM-VLM-Base, Qwen2.5-VL based) on robot manipulation data to turn it into an 'embodied brain'; variants fine-tuned on Unitree open datasets and LIBERO. (https://huggingface.co/collections/unitreerobotics/unifolm-vla-0)
  - World-model-action architecture (WMA-0): UnifoLM-WMA-0 (Sept 2025, Apache-2.0) pairs a world model that predicts future interactions (usable as a simulator) with action generation; Base and Dual variants on HF. (https://huggingface.co/unitreerobotics/UnifoLM-WMA-0-Base)


## VUI Labs

### VUI Labs Luna-TTS (and Luna-TTS Realtime)
**VUI Labs Luna-TTS (and Luna-TTS Realtime)** (VUI Labs; current; audio/speech; released 2026-06) | VUI Labs API: https://www.vuilabs.ai/; arXiv (technical report): https://arxiv.org/abs/2608.11593 — Chinese voice-AI startup (Pandaily). Release month June 2026 per the Artificial Analysis leaderboard; technical report 2026-08-12 (Feng Yin et al., 22 authors). Supports zero-shot cloning, speech editing, emotion control, non-verbal vocalisations. We found no statement about open weights. Pandaily headline calls it China's 'Thinking Machines' and names Qian Yanmin (role not verified). Not the same as fluxions-ai 'Vui' (open Apache-2.0 small TTS).
  - VUI Labs API — https://www.vuilabs.ai/
  - arXiv (technical report) — https://arxiv.org/abs/2608.11593
  - Pricing source: https://www.vuilabs.ai/
Notable capabilities:
  - Diffusion-language-model TTS (non-autoregressive): Generates the whole RVQ token grid in a fixed number of parallel refinement steps; the Realtime variant is blockwise-autoregressive over 1.28 s blocks (RTF 0.0240, 41.6 ms first-block latency locally). 0.6B backbone, ~1M hours of zh/en/ja/ko speech. (https://arxiv.org/abs/2608.11593)
  - Chinese startup at the top of TTS arenas (found after launch): Pandaily (Aug 2026) reported #1 on Hugging Face TTS Arena and #3 on Artificial Analysis Speech Arena; on 2026-09-29 AA shows it #8 (Elo 1230). (https://pandaily.com/vui-labs-luna-tts-number-one-tts-arena-qian-yanmin-voice-agent-aug2026)


## xAI

### Grok 4.7
**Grok 4.7** (xAI; current; reasoning-llm; released 2026-09-21) | ctx 500,000 | $2 in / $6 out per 1M tokens (USD); higher tier applies to whole request when prompt >= 200k tokens | xAI API: `grok-4.7`; AWS Bedrock: `xai.grok-4.7`; OpenRouter: `x-ai/grok-4.7` | Web app: https://grok.com — Alias grok-4.7-latest. xAI flagship as of Sept 2026; no Batch API; logprobs unsupported. Bedrock launched 2026-09-28 (Global CRIS $2/$6, Geo $2.20/$6.60). Max output not published.
  - xAI API: `grok-4.7` — https://api.x.ai/v1/chat/completions (docs: https://docs.x.ai/docs/models/grok-4.7)
  - AWS Bedrock: `xai.grok-4.7` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-7.html)
  - OpenRouter: `x-ai/grok-4.7` — https://openrouter.ai/x-ai/grok-4.7
  - Web app — https://grok.com
  - Pricing source: https://docs.x.ai/docs/models
Notable capabilities:
  - Four-level reasoning effort incl. xhigh: Configurable reasoning effort low / medium / high / xhigh (default high) on one model id. (https://docs.x.ai/docs/models/grok-4.7)
  - 500K context at unchanged price: 500K-token context with text+image input; launched at the same $2/$6 price as Grok 4.6 while claiming notable gains. (https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-7.html)
  - Mixed independent benchmark results (found after launch): Early third-party evals showed a more mixed picture than xAI's claims, still behind top Claude/GPT-6 models on several tasks. (https://tech.yahoo.com/ai/gemini/articles/xai-launches-grok-4-7-171603280.html)

### Grok Voice Transcribe 2.0
**Grok Voice Transcribe 2.0** (xAI; current; audio/speech; released 2026-09-18) | xAI API (REST): `grok-voice-transcribe-2.0`; xAI API (WebSocket streaming): `grok-voice-transcribe-2.0` — Launched 2026-09-18 as drop-in upgrade of the Grok STT API (first released 2026-04-17 with grok-voice-transcribe-1.0, which can be pinned but will be deprecated). Up to 500 MB files; WAV/MP3/OGG/Opus/FLAC/AAC/MP4/M4A/MKV plus raw PCM/mu-law/A-law at 8-48 kHz; Smart Turn end-of-turn detection, VAD, inverse text normalization, filler removal, mid-recording language switching. Docs list ~25 languages for formatting.
  - xAI API (REST): `grok-voice-transcribe-2.0` — https://api.x.ai/v1/stt (docs: https://docs.x.ai/developers/model-capabilities/audio/speech-to-text)
  - xAI API (WebSocket streaming): `grok-voice-transcribe-2.0` — wss://api.x.ai/v1/stt (docs: https://docs.x.ai/developers/model-capabilities/audio/speech-to-text)
  - Pricing source: https://x.ai/news/grok-voice-transcribe-2
Notable capabilities:
  - Top streaming STT accuracy (claimed): xAI says it ranks #1 for accuracy among 32 streaming models on the Artificial Analysis leaderboard; multilingual short-phrase WER 20.6% -> 6.8% vs v1.0 ('2x as accurate'). (https://x.ai/news/grok-voice-transcribe-2)
  - Very low price with diarization included: $0.10/hr batch and $0.20/hr streaming, with speaker diarization, word timestamps, up to 8-channel multichannel and 100 key terms per request at no extra cost. (https://x.ai/news/grok-voice-transcribe-2)

### Grok Imagine Image 2.0
**Grok Imagine Image 2.0** (xAI; current; image-gen; released 2026-08-07) | xAI API: `grok-imagine-image-2.0`; xAI API (edits): `grok-imagine-image-2.0` | Web app: https://grok.com — xAI's recommended image model; cheaper grok-imagine-image ($0.02) and grok-imagine-image-quality ($0.05) also listed. App launch 2026-08-07, API shortly after. Third-party reports of resolution/quality price tiers not verified on official page.
  - xAI API: `grok-imagine-image-2.0` — https://api.x.ai/v1/images/generations (docs: https://docs.x.ai/docs/guides/image-generation)
  - xAI API (edits): `grok-imagine-image-2.0` — https://api.x.ai/v1/images/edits
  - Web app — https://grok.com
  - Pricing source: https://docs.x.ai/docs/models
Notable capabilities:
  - Generation + editing in one model: Text-to-image and image editing (URL or base64 input) via /v1/images/generations and /v1/images/edits. (https://docs.x.ai/docs/guides/image-generation)
  - Top-2 on Arena image leaderboards at launch: xAI reported #2 on both Arena Text-to-Image and Arena Image Edit at launch (Aug 7, 2026). (https://kie.ai/blog/grok-imagine-image-2-0-release)

### Grok Voice Think Fast 2.0
**Grok Voice Think Fast 2.0** (xAI; current; audio/speech; released 2026-07-29) | xAI API (Voice Agent / speech-to-speech, WebSocket): `grok-voice-think-fast-2.0`; xAI API (alias): `grok-voice-latest` | Web app: https://grok.com — Released 2026-07-29; grok-voice-latest switched to it on 2026-08-05. Predecessor grok-voice-think-fast-1.0 can still be pinned. 20+ languages; audio PCM (8-48 kHz), Opus 24 kHz, G.711 mu-law/A-law; server VAD, session resumption (30 min), custom cloned voices. xAI says Starlink A/B tests raised sales conversion and support containment. Benchmarks are xAI-reported.
  - xAI API (Voice Agent / speech-to-speech, WebSocket): `grok-voice-think-fast-2.0` — wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0 (docs: https://docs.x.ai/developers/model-capabilities/audio/voice-agent)
  - xAI API (alias): `grok-voice-latest` (docs: https://docs.x.ai/developers/model-capabilities/audio/voice-agent)
  - Web app — https://grok.com
  - Pricing source: https://docs.x.ai/developers/pricing
Notable capabilities:
  - Reasoning while speaking: Speech-to-speech model that reasons in real time (reasoning effort 'high' by default, can be set to 'none'); 97.2% Big Bench Audio, 82.9 on the Artificial Analysis Speech-to-Speech Quality Index (vs 75.7 for v1.0). (https://x.ai/news/grok-voice-think-fast-2)
  - Faster first audio: Time to first audio cut from 1.25 s (v1.0) to 0.70 s; Full Duplex Bench 95.1%, tau-voice Bench 56.5% (xAI-reported). (https://x.ai/news/grok-voice-think-fast-2)
  - Built-in server-side tools: Web search, X search, collections (file) search and remote MCP callable from inside a voice session, plus custom functions. (https://docs.x.ai/developers/model-capabilities/audio/voice-agent)
  - OpenAI Realtime-compatible protocol: Largely compatible with the OpenAI Realtime SDK: change base URL to https://api.x.ai/v1 and the API key (minor event-name differences). (https://docs.x.ai/developers/model-capabilities/audio/voice-agent)

### Grok Imagine Video 1.5
**Grok Imagine Video 1.5** (xAI; current; video-gen; released 2026-05-30) | xAI API: `grok-imagine-video-1.5` | Web app: https://grok.com — Snapshot alias grok-imagine-video-1.5-2026-05-30. Legacy grok-imagine-video still available at $0.05/s. Resolution/audio details not verified.
  - xAI API: `grok-imagine-video-1.5` — https://api.x.ai/v1/videos/generations (docs: https://docs.x.ai/docs/guides/video-generation)
  - Web app — https://grok.com
  - Pricing source: https://docs.x.ai/docs/models
Notable capabilities:
  - Image-to-video up to 15 s: Animates a source still (URL/base64) or prompt into clips up to 15 seconds; async job polled via GET /v1/videos/{request_id}. (https://docs.x.ai/docs/guides/video-generation)
  - Per-second pricing, text or image input: Text- or image-to-video at $0.08 per generated second (legacy grok-imagine-video $0.05/s). (https://docs.x.ai/docs/models)

### Grok Build 0.1
**Grok Build 0.1** (xAI; current; code; released 2026-05) | ctx 256,000 | $1 in / $2 out per 1M tokens (USD) | xAI API: `grok-build-0.1`; OpenRouter: `x-ai/grok-build-0.1` — xAI coding model (successor to grok-code-fast line). Release month from OpenRouter listing (2026-05-20); exact date not verified.
  - xAI API: `grok-build-0.1` — https://api.x.ai/v1/chat/completions (docs: https://docs.x.ai/docs/models/grok-build-0.1)
  - OpenRouter: `x-ai/grok-build-0.1` — https://openrouter.ai/x-ai/grok-build-0.1
  - Pricing source: https://docs.x.ai/docs/models/grok-build-0.1
Notable capabilities:
  - Agentic coding model: Reasoning model tuned for agentic software engineering and workflow tasks; powers xAI's Grok Build coding agent. (https://docs.x.ai/docs/models/grok-build-0.1)
  - Low-cost coding tier: $1/$2 per 1M tokens with 256K context - cheapest current Grok text model. (https://docs.x.ai/docs/models)

### Grok Text to Speech (Grok TTS API)
**Grok Text to Speech (Grok TTS API)** (xAI; current; audio/speech; released 2026-04-17) — Launched with the Grok STT API on 2026-04-17 (some press reports an earlier developer opening in March 2026). No separate model id is documented; the endpoint selects the model. 60,000 characters per REST request; ~20 languages plus auto-detect; MP3/WAV/PCM/mu-law/A-law at 8-48 kHz; voice list via GET /v1/tts/voices (Ara, Eve, Leo, Rex, Sal and many more).
  - xAI API (REST) — https://api.x.ai/v1/tts (docs: https://docs.x.ai/developers/model-capabilities/audio/text-to-speech)
  - xAI API (WebSocket streaming) — wss://api.x.ai/v1/tts (docs: https://docs.x.ai/developers/model-capabilities/audio/text-to-speech)
  - Pricing source: https://docs.x.ai/developers/pricing
Notable capabilities:
  - Inline speech tags: Inline tags ([pause], [laugh], [sigh], [cry], [gasp], ...) and wrapping tags (<whisper>, <soft>, <loud>, <slow>, <fast>, <sing>) control delivery. (https://docs.x.ai/developers/model-capabilities/audio/text-to-speech)
  - Custom (cloned) voices (found after launch): Clone a voice from a short reference clip via the Custom Voices API; the voice_id works like built-in voices in TTS and the Voice Agent API. (https://docs.x.ai/developers/model-capabilities/audio/text-to-speech)

### Grok 4.3
**Grok 4.3** (xAI; current; reasoning-llm; released 2026-04) | ctx 1,000,000 | $1.25 in / $2.5 out per 1M tokens (USD); Batch API 20% off | xAI API: `grok-4.3`; AWS Bedrock: `xai.grok-4.3`; OpenRouter: `x-ai/grok-4.3` | Web app: https://grok.com — Alias grok-4.3-latest. Cheaper long-context option still offered alongside Grok 4.7. Release month inferred from OpenRouter listing date (2026-04-30); exact date not verified.
  - xAI API: `grok-4.3` — https://api.x.ai/v1/chat/completions (docs: https://docs.x.ai/docs/models/grok-4.3)
  - AWS Bedrock: `xai.grok-4.3` (docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html)
  - OpenRouter: `x-ai/grok-4.3` — https://openrouter.ai/x-ai/grok-4.3
  - Web app — https://grok.com
  - Pricing source: https://docs.x.ai/docs/models
Notable capabilities:
  - 1M context at budget price: 1M-token context window at $1.25/$2.50, cheaper than the 500K-context Grok 4.5-4.7 line. (https://docs.x.ai/docs/models/grok-4.3)
  - Reasoning effort incl. none: Reasoning effort none / low / medium / high / xhigh, default low - usable as a fast non-reasoning model. (https://docs.x.ai/docs/models/grok-4.3)

### Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)
**Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)** (xAI; legacy; reasoning-llm; released 2026-03) | ctx 1,000,000 | $1.25 in / $2.5 out per 1M tokens (USD) | xAI API: `grok-4.20-0309-reasoning`; xAI API (non-reasoning): `grok-4.20-0309-non-reasoning`; xAI API (multi-agent): `grok-4.20-multi-agent-0309`; OpenRouter: `x-ai/grok-4.20`; OpenRouter (multi-agent): `x-ai/grok-4.20-multi-agent` — Snapshot ids dated 0309. xAI docs list 1M context; OpenRouter lists 2M. Superseded by Grok 4.5-4.7; logprobs unsupported.
  - xAI API: `grok-4.20-0309-reasoning` — https://api.x.ai/v1/chat/completions (docs: https://docs.x.ai/docs/models)
  - xAI API (non-reasoning): `grok-4.20-0309-non-reasoning` — https://api.x.ai/v1/chat/completions
  - xAI API (multi-agent): `grok-4.20-multi-agent-0309` (docs: https://docs.x.ai/developers/model-capabilities/text/multi-agent)
  - OpenRouter: `x-ai/grok-4.20` — https://openrouter.ai/x-ai/grok-4.20
  - OpenRouter (multi-agent): `x-ai/grok-4.20-multi-agent` — https://openrouter.ai/x-ai/grok-4.20-multi-agent
  - Pricing source: https://docs.x.ai/docs/models
Notable capabilities:
  - Multi-agent model variant: Dedicated API id that runs parallel collaborating agents (4 at low/medium effort, 16 at high/xhigh) that search and cross-check before synthesizing an answer. (https://docs.x.ai/developers/model-capabilities/text/multi-agent)
  - Reasoning and non-reasoning twin ids: Same snapshot (0309) offered as separate reasoning and non-reasoning model ids. (https://docs.x.ai/docs/models)


## Xiaomi

### Xiaomi-Robotics-1 (XR-1, 5B)
**Xiaomi-Robotics-1 (XR-1, 5B)** (Xiaomi; current; robotics; released 2026-07-16; open weights) | Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-1-5B` | GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-1 — Paper 2026-07-16 (arXiv 2607.15330); weights on HF 2026-07-28; code 2026-08-03. Predecessor: xiaomi-robotics-0 (Feb 2026, arXiv 2602.12684). Companion world model: xiaomi-robotics-u0 (July/Sept 2026). Changelog 2026-09-29: linked the new model files.
  - Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-1-5B` — https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-1-5B
  - GitHub — https://github.com/XiaomiRobotics/Xiaomi-Robotics-1 (docs: https://robotics.xiaomi.com/robot-static-resource/xiaomi-robotics-1/xiaomi-robotics-1.pdf)
Notable capabilities:
  - VLA pretrained on 100K+ hours of real trajectories: Pretrained on 100K+ hours of embodiment-free UMI trajectories across 1,700+ scenarios (per Xiaomi project materials), then post-trained on 10K+ hours of cross-embodiment data, for out-of-the-box mobile manipulation in unseen environments. (https://arxiv.org/abs/2607.15330)
  - Open-weight SOTA on sim benchmarks: RoboCasa 74.5%, RoboCasa365 57.4%, VLABench 59.1%, RoboDojo 13.93% in the GitHub table (the arXiv abstract cites a 20.07 RoboDojo average score — different metric/version), each ahead of the runner-up per the authors. (https://github.com/XiaomiRobotics/Xiaomi-Robotics-1)

### Xiaomi-Robotics-U0 (38B) / U0-4B
**Xiaomi-Robotics-U0 (38B) / U0-4B** (Xiaomi; current; world-model; released 2026-07-13; open weights) | Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-U0`; Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-U0-4B` | GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0; ModelScope: https://modelscope.cn/collections/XiaomiRobotics/Xiaomi-Robotics-U0 — 'first' flag is Xiaomi's claim (first model with high-quality multi-view scene generation across multiple robot embodiments). Paper says 38B params; the HF README table says 34B. U0 and U0-FlashAR weights 2026-07-13; U0-4B, U0-Sequence and U0-4B-Sequence weights plus FSDP training code 2026-09-08. U0-Video announced as coming soon. Authors report beating GPT-Image-2.0 in human evals of embodied scene generation/transfer and #1 on World Arena for embodied video. Not an action model: it generates observations/data, not motor commands.
  - Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-U0` — https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0
  - Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-U0-4B` — https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B
  - GitHub — https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0
  - ModelScope — https://modelscope.cn/collections/XiaomiRobotics/Xiaomi-Robotics-U0
Notable capabilities:
  - [FIRST] Unified embodied synthesis: One autoregressive model (shared discrete visual tokenizer, next-token objective, initialized from Emu3.5) does text-to-image, image editing, multi-view robot scene generation, embodied transfer (editing scenes while keeping multi-view consistency) and embodied video rollout. (https://arxiv.org/abs/2607.11643)
  - Data engine for VLAs: Synthetic data from U0 raised π0.5's out-of-distribution success on hard real-world manipulation tasks from 36.9% to 63.2% (authors). (https://arxiv.org/abs/2607.11643)
  - FlashAR fast decoding: Anti-diagonal grouped visual-token decoding plus vLLM batching: 5.44 s per 1024x1024 image on one H20, 82.86x faster than eager AR. (https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B)

### Xiaomi-Robotics-0 (4.7B VLA)
**Xiaomi-Robotics-0 (4.7B VLA)** (Xiaomi; legacy; robotics; released 2026-02-12; open weights) | Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-0-Pretrain` | Project page: https://xiaomi-robotics-0.github.io — 4.7B parameters, Qwen3-VL-4B-Instruct backbone; pretrained on cross-embodiment robot trajectories plus vision-language data. Real-robot evals: Lego disassembly and towel folding (bimanual). Checkpoints: -Pretrain, -LIBERO, -Calvin-ABC_D, -Calvin-ABCD_D, -SimplerEnv-WidowX, -SimplerEnv-Google-Robot (HF, 2026-02-10). Paper arXiv 2602.12684 (2026-02-13). Superseded by xiaomi-robotics-1 (July 2026).
  - Hugging Face: `XiaomiRobotics/Xiaomi-Robotics-0-Pretrain` — https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-0-Pretrain
  - Project page — https://xiaomi-robotics-0.github.io
Notable capabilities:
  - Real-time asynchronous execution on a consumer GPU: Post-trained for asynchronous execution with aligned timesteps between consecutive action chunks, so rollouts stay smooth despite inference latency; runs on a consumer-grade GPU (per paper). (https://arxiv.org/abs/2602.12684)
  - Strong open sim-benchmark results: LIBERO 98.7% avg; SimplerEnv Visual Matching 85.5%, Visual Aggregation 74.7%, WidowX 79.2%; CALVIN avg length 4.75 (ABC-D) / 4.80 (ABCD-D) (authors). (https://xiaomi-robotics-0.github.io)


## Zhipu AI (Z.ai)

### GLM-5.3-Flash / FlashX
**GLM-5.3-Flash / FlashX** (Zhipu AI (Z.ai); current; multimodal; released 2026-08; open weights) | ctx 1,000,000 | $0.15 in / $0.5 out per 1M tokens (USD) for glm-5.3-flash; glm-5.3-flashx (~200 tok/s): 0.37 in / 1.25 out / 0.075 cached | Z.ai API: `glm-5.3-flash`; Z.ai API (fast): `glm-5.3-flashx`; OpenRouter: `z-ai/glm-5.3-flash`; OpenRouter (FlashX): `z-ai/glm-5.3-flashx` | Hugging Face: https://huggingface.co/zai-org/GLM-5.3-Flash; Web app: https://chat.z.ai — Z.ai says it beats GLM-5.2 at a fraction of the cost; 3x Coding Plan quota vs GLM-5.3 (FlashX not yet on the plan). Thinking cannot be disabled. 'first' claim is the vendor's own.
  - Z.ai API: `glm-5.3-flash` — https://api.z.ai/api/paas/v4 (docs: https://docs.z.ai/guides/vlm/glm-5.3-flash)
  - Z.ai API (fast): `glm-5.3-flashx` — https://api.z.ai/api/paas/v4 (docs: https://docs.z.ai/guides/vlm/glm-5.3-flash)
  - OpenRouter: `z-ai/glm-5.3-flash` — https://openrouter.ai/z-ai/glm-5.3-flash
  - OpenRouter (FlashX): `z-ai/glm-5.3-flashx` — https://openrouter.ai/z-ai/glm-5.3-flashx
  - Hugging Face — https://huggingface.co/zai-org/GLM-5.3-Flash
  - Web app — https://chat.z.ai
  - Pricing source: https://docs.z.ai/guides/overview/pricing
Notable capabilities:
  - First native multimodal GLM-5 model: First GLM-5-series model with native vision (image, video, file input); vision used inside the coding loop (UI replication, Blender, browser/computer-use agents). (https://docs.z.ai/guides/vlm/glm-5.3-flash)
  - [FIRST] Sparse + linear attention hybrid: 320B total / 18B active; Z.ai claims it is the first open-source frontier model combining sparse and linear attention (3.01x less attention compute, 4.44x smaller KV cache vs GLM-5.3). (https://docs.z.ai/guides/vlm/glm-5.3-flash)
  - Office deliverables with visual self-check: Produces PPTX/PDF/DOCX/XLSX and renders them to catch overflow and layout issues. (https://docs.z.ai/guides/vlm/glm-5.3-flash)

### GLM-5.3
**GLM-5.3** (Zhipu AI (Z.ai); current; reasoning-llm; released 2026-08; open weights) | ctx 1,000,000 | $1.4 in / $4.4 out per 1M tokens (USD) | Z.ai API: `glm-5.3`; Z.ai API (Anthropic format): `glm-5.3`; Alibaba Cloud Model Studio: `ZHIPU/GLM-5.3`; OpenRouter: `z-ai/glm-5.3` | Hugging Face: https://huggingface.co/zai-org/GLM-5.3; Web app: https://chat.z.ai — Z.ai flagship. Text-only input. Migration: requests with thinking disabled fail - set enabled + reasoning_effort low. Coding Plan base URL is https://api.z.ai/api/coding/paas/v4. Release day not verified (OpenRouter 2026-08-18, HF 2026-08-25). GLM-5.2 (same price, MIT weights) still listed.
  - Z.ai API: `glm-5.3` — https://api.z.ai/api/paas/v4 (docs: https://docs.z.ai/guides/llm/glm-5.3)
  - Z.ai API (Anthropic format): `glm-5.3` — https://api.z.ai/api/anthropic (docs: https://docs.z.ai/guides/llm/glm-5.3)
  - Alibaba Cloud Model Studio: `ZHIPU/GLM-5.3` (docs: https://www.alibabacloud.com/help/en/model-studio/models)
  - OpenRouter: `z-ai/glm-5.3` — https://openrouter.ai/z-ai/glm-5.3
  - Hugging Face — https://huggingface.co/zai-org/GLM-5.3
  - Web app — https://chat.z.ai
  - Pricing source: https://docs.z.ai/guides/overview/pricing
Notable capabilities:
  - Post-training-only jump in coding: Same base as GLM-5.2; Z.ai reports +50% on its Code Bench and open-model SOTA on Terminal Bench 3.0 and Agents' Last Exam (CLI). (https://docs.z.ai/guides/llm/glm-5.3)
  - Emergent cyber capability: Best CyberGym vulnerability-discovery score to date per Z.ai; exploitation benchmark scores more than double GLM-5.2's. (https://docs.z.ai/guides/llm/glm-5.3)
  - Always-on reasoning with effort levels: thinking.type disabled no longer allowed; reasoning_effort low/high/max (default max). (https://docs.z.ai/guides/llm/glm-5.3)
  - Coding Plan integration: Available in the GLM Coding Plan (points-based; off-peak/weekend calls cost 50% points) for Claude Code, Cline, OpenCode etc. (https://docs.z.ai/guides/llm/glm-5.3)

### GLM-4.6V
**GLM-4.6V** (Zhipu AI (Z.ai); legacy; multimodal; released 2025-12; open weights) | ctx 128,000 | $0.3 in / $0.9 out per 1M tokens (USD); GLM-4.6V-FlashX 0.04/0.4; GLM-4.6V-Flash free | Z.ai API: `glm-4.6v`; OpenRouter: `z-ai/glm-4.6v` | Hugging Face: https://huggingface.co/zai-org/GLM-4.6V; Hugging Face (Flash): https://huggingface.co/zai-org/GLM-4.6V-Flash; Web app: https://chat.z.ai — Still sold on Z.ai (with FlashX and free Flash variants) but superseded by the natively multimodal GLM-5.3-Flash. Model id casing on Z.ai assumed lowercase glm-4.6v (listed as GLM-4.6V). Release day not verified (HF 2025-12-07).
  - Z.ai API: `glm-4.6v` — https://api.z.ai/api/paas/v4 (docs: https://docs.z.ai/guides/overview/pricing)
  - OpenRouter: `z-ai/glm-4.6v` — https://openrouter.ai/z-ai/glm-4.6v
  - Hugging Face — https://huggingface.co/zai-org/GLM-4.6V
  - Hugging Face (Flash) — https://huggingface.co/zai-org/GLM-4.6V-Flash
  - Web app — https://chat.z.ai
  - Pricing source: https://docs.z.ai/guides/overview/pricing
Notable capabilities:
  - Native multimodal function calling: First GLM vision model with native function calling (images can be passed to and returned from tools). (https://huggingface.co/zai-org/GLM-4.6V)
  - Interleaved image-text generation: Builds mixed image-text content from documents and tool-retrieved images; also frontend replication from screenshots. (https://huggingface.co/zai-org/GLM-4.6V)

