{"schema":"postcutoff/itemlist@1","as_of":"2026-10-07T23:43:00+02:00","url":"https://postcutoff.com/models/","md":"https://postcutoff.com/models/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"page":1,"pages":1,"total":298,"items":[{"id":"embeddinggemma-2","name":"EmbeddingGemma 2","org":"Google DeepMind","family":"Gemma","released":"2026-10-06","status":"current","type":"embedding","modality_in":["text","image","video","audio"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":8192,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"google/embeddinggemma-2","url":"https://huggingface.co/google/embeddinggemma-2"},{"provider":"Google docs","docs":"https://ai.google.dev/gemma/docs/embeddinggemma"}],"capabilities":[{"name":"Four modalities in one small open embedding space","detail":"Maps text (including code), images, video and audio into one 768-dimensional space with 740M parameters in total: a 270M text model (130M backbone + 140M embedder) plus optional 170M vision and 300M audio encoders that can be loaded separately.","first":false,"discovered":"launch","source":"https://huggingface.co/google/embeddinggemma-2"},{"name":"Matryoshka truncation","detail":"Embeddings can be cut to 512, 256 or 128 dimensions (up to 6x less storage); quality loss is small down to 256d.","first":false,"discovered":"launch","source":"https://huggingface.co/google/embeddinggemma-2"},{"name":"Big code-retrieval gain over v1","detail":"MTEB code v1 78.68 vs 68.76 for EmbeddingGemma 1; MTEB multilingual v2 61.36 vs 61.15; MMEB v2 overall 59.01 (Google's numbers).","first":false,"discovered":"launch","source":"https://huggingface.co/google/embeddinggemma-2"}],"entry":"2026-10-06-embeddinggemma-2","notes":"Announced 2026-10-06 on the Google blog. Apache 2.0 (Gemma 4 license), not gated on Hugging Face. Runs on phones and laptops; uses task-instruction prefixes (e.g. 'task: search result | query: ...') for text inputs. sentence-transformers id: google/embeddinggemma-2. No hosted Gemini API endpoint is listed; for hosted multimodal embeddings Google offers gemini-embedding-2.","verified":"2026-10-07","body_md":"Small open on-device embedding model for multimodal search, RAG, classification and clustering.\n\n```python\nfrom sentence_transformers import SentenceTransformer\nmodel = SentenceTransformer(\"google/embeddinggemma-2\")\nq = model.encode(\"What causes the northern lights?\", prompt_name=\"SearchQuery\")\nd = model.encode(\"The northern lights are caused by charged particles from the sun.\", prompt_name=\"Document\")\nprint(model.similarity(q, d))\n```\n\nSources: [Hugging Face model card](https://huggingface.co/google/embeddinggemma-2), [launch blog](https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/), [docs](https://ai.google.dev/gemma/docs/embeddinggemma).","page_url":"https://postcutoff.com/m/embeddinggemma-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":6,"you_url":null,"briefings":null},{"id":"mistral-large-4","name":"Mistral Large 4","org":"Mistral AI","family":"Mistral Large","released":"2026-10-06","status":"preview","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"not yet stated (weights promised by end of Oct 2026)","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.68,"output":2.09,"unit":"per 1M tokens (USD); docs price, blog lists $1.36/$4.18 - unverified which applies after preview","source":"https://docs.mistral.ai/models/mistral-large-4-0"},"price_line":"$0.68 in, $2.09 out per 1M tokens","access":[{"provider":"Mistral API","model_id":"mistral-large-4-0","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://docs.mistral.ai/models/mistral-large-4-0"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"Trillion-parameter open-weight European model","detail":"Granular MoE with 49B active / 1.05T total parameters and 1M context; weights promised by end of October 2026.","first":true,"discovered":"launch","source":"https://mistral.ai/news/mistral-large-4/"},{"name":"Cybersecurity","detail":"Mistral reports Cybench 93% and a top-5 place on the Artificial Analysis Cyber Index.","first":false,"discovered":"launch","source":"https://mistral.ai/news/mistral-large-4/"}],"entry":"2026-10-06-mistral-large-4","notes":"Public preview from Oct 6, 2026. Prices conflict between docs and blog.","verified":"2026-10-06","body_md":"Mistral's flagship open-weight MoE, successor to Mistral Large 3.","page_url":"https://postcutoff.com/m/mistral-large-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":6,"you_url":"https://postcutoff.com/you/mistral-large-4/","briefings":null},{"id":"gemini-nano-banana-2-1","name":"Nano Banana 2.1 (Gemini Nano Banana 2.1)","org":"Google DeepMind","family":"Gemini Image","released":"2026-10-06","status":"current","type":"image-gen","modality_in":["text","image","video","pdf"],"modality_out":["image","text"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":32768,"knowledge_cutoff":"2026-03","pricing":{"input":1.5,"output":7.5,"per_image_1k":0.0336,"per_image_2k":0.0504,"per_image_4k":0.113,"unit":"per 1M tokens: $1.50 input (text/image/video), $7.50 text+thinking output, $30 image output (= $0.0336 per 1K, $0.0504 per 2K, $0.113 per 4K image); Batch is half price; no free tier","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1.50 in, $7.50 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-nano-banana-2.1","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-nano-banana-2.1:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-nano-banana-2.1"},{"provider":"Google AI Studio","url":"https://aistudio.google.com?model=gemini-nano-banana-2.1"},{"provider":"OpenRouter","model_id":"google/gemini-nano-banana-2.1"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Half the per-image price of Nano Banana 2","detail":"Image output is $30 per 1M tokens versus $60 for gemini-3.1-flash-image: a 1K image costs $0.0336 instead of $0.067. Input is pricier ($1.50 vs $0.50 per 1M).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/pricing"},{"name":"Built on Gemini 3.6 Flash","detail":"Model card: 'Nano Banana 2.1 is based on Gemini 3.6 Flash', with a March 2026 knowledge cutoff.","first":false,"discovered":"launch","source":"https://deepmind.google/models/model-cards/nano-banana-2-1/"},{"name":"Configurable thinking and search grounding","detail":"Thinking levels minimal / medium (default) / high; grounding with Google Web and Image Search; up to 14 reference images (4 consistent characters, 10 objects).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-nano-banana-2.1"},{"name":"Fixed panorama tiling","detail":"Fixes tiling artifacts at 1:4, 4:1, 1:8 and 8:1 aspect ratios at 2K and 4K, and improves text and infographic layout.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-nano-banana-2.1"}],"entry":"2026-10-06-nano-banana-2-1","notes":"GA on the Gemini API on 2026-10-06 (stable id only, no preview id). Also rolled out in the Gemini app, AI Mode in Search, Flow, Stitch, Google Ads and Gemini Enterprise (per The Decoder). It deprecates gemini-3.1-flash-image (no shutdown date yet). No caching, function calling or structured outputs. The Decoder quotes a $0.0756 4K price; the official pricing page says $0.113 (3,780 tokens), which is the figure used here.","verified":"2026-10-07","body_md":"Google's default fast image generation/editing model from Oct 2026, successor to Nano Banana 2 (gemini-3.1-flash-image).\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-nano-banana-2.1:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"An infographic about the water cycle\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-nano-banana-2.1), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [Gemini API changelog, Oct 6 2026](https://ai.google.dev/gemini-api/docs/changelog), [model card](https://deepmind.google/models/model-cards/nano-banana-2-1/), [The Decoder](https://the-decoder.com/googles-new-image-model-nano-banana-2-1-generates-better-images-for-less-money/).","page_url":"https://postcutoff.com/m/gemini-nano-banana-2-1/","events_after":747,"major_after":187,"historic_after":40,"missing_at_launch":706,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-03-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-03-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-03-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-03.md"}},{"id":"reflection-beam","name":"Beam (Beam-501B-A23B)","org":"Reflection AI","family":"Beam","released":"2026-10-05","status":"preview","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":256000,"max_output":128000,"knowledge_cutoff":"2026-06","pricing":null,"price_line":"Price not published","access":[{"provider":"Reflection API (beta, early access / waitlist)","model_id":"Beam-501B-A23B","endpoint":"https://api.reflection.ai/v1","docs":"https://developers.reflection.ai/models"},{"provider":"Reflection API, OpenAI-compatible (Chat Completions, Models)","model_id":"Beam-501B-A23B","endpoint":"https://api.reflection.ai/openai/v1","docs":"https://developers.reflection.ai/"},{"provider":"Web app / early-access signup","url":"https://platform.reflection.ai/"}],"capabilities":[{"name":"Efficient 500B-class open MoE","detail":"501B total / 23B active sparse MoE trained on 23.8T tokens; Reflection says it matches GLM-5.2 on reasoning with 3-4x less inference compute.","first":false,"discovered":"launch","source":"https://reflection.ai/blog/introducing-beam"},{"name":"Large-scale agentic RL","detail":"RL on 10.5K GB300 GPUs for 4 weeks: >100M rollouts across ~1M coding, agentic and STEM environments using ~1.3B sandboxes.","first":false,"discovered":"launch","source":"https://reflection.ai/blog/introducing-beam"},{"name":"1M-token context (training)","detail":"Pretrained at 256K and extended to 1M tokens in midtraining; the beta API currently exposes 256K context and 128K output.","first":false,"discovered":"launch","source":"https://developers.reflection.ai/models"}],"entry":"2026-10-05-reflection-beam-501b-open-weight","notes":"Announced Oct 5, 2026; Apache-2.0 weights and tech report promised for later in October 2026 (no Hugging Face repo yet as of Oct 5). API is beta, reasoning always on (default effort medium), tool calling and structured outputs. No published price. Benchmarks are Reflection's own: SWE-bench Verified 80.9, Terminal Bench v2.1 80.1, GPQA Diamond 90.5, HLE no tools 36.2.","verified":"2026-10-05","body_md":"Reflection AI's first model. Until the weights ship, access is through the waitlisted beta API. The OpenAI-compatible endpoint\nworks with the official OpenAI SDKs once you change `base_url` to `https://api.reflection.ai/openai/v1` and use a Reflection API key.\nWhen the weights are released, add the Hugging Face repo and switch `status` to `current`.","page_url":"https://postcutoff.com/m/reflection-beam/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":589,"events_since_release":null,"you_url":"https://postcutoff.com/you/reflection-beam/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"kolibri-1","name":"Kolibri-1","org":"Aleph Alpha","family":"Kolibri","released":"2026-10-03","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":1048576,"max_output":null,"knowledge_cutoff":"2026-06","pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/Aleph-Alpha/Kolibri-1"},{"provider":"Self-hosted (vLLM + aleph-alpha-inference, OpenAI-compatible API)","docs":"https://huggingface.co/Aleph-Alpha/Kolibri-1"}],"capabilities":[{"name":"German-optimised tokenizer","detail":"Bilingual EN–DE 128k tokenizer; Aleph Alpha reports 11.2% fewer German tokens than GPT-5's tokenizer (an independent test on the German Basic Law found ~15%).","first":false,"discovered":"launch","source":"https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/"},{"name":"Sparse 1M-token context","detail":"78B-total / ~3.46B-active MoE (384 experts, 6 routed) with 512-token sliding-window attention and a global layer every 5th layer; up to 1,048,576 tokens of context (262,144 recommended).","first":false,"discovered":"launch","source":"https://huggingface.co/Aleph-Alpha/Kolibri-1"},{"name":"Abstention / grounding training","detail":"Trained with a Merlin–Arthur protocol to abstain rather than hallucinate; Aleph Alpha reports an M/A grounding score of 0.23 (scale 0–0.5).","first":false,"discovered":"launch","source":"https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/"}],"entry":"2026-10-03-aleph-alpha-kolibri-open-weights","notes":"Reasoning mode with configurable effort and tool calling. Recommended sampling temperature 1.0, top_p 0.97, top_k 128. Needs ~2×80GB GPUs (min 2× A100 80GB). Benchmarks are Aleph Alpha's own (AIME 2025 96.9 EN, GPQA Diamond 84.3 EN). Weaker than Qwen models on closed-book recall and agentic coding per reviewers. No hosted API price found.","verified":"2026-10-03","body_md":"Aleph Alpha's open-weight English–German MoE. Serve with vLLM plus the `aleph-alpha-inference` package (reasoning and tool-calling parsers enabled) for an OpenAI-compatible endpoint.","page_url":"https://postcutoff.com/m/kolibri-1/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":569,"events_since_release":null,"you_url":"https://postcutoff.com/you/kolibri-1/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"cactus-whistle","name":"Cactus Whistle (16.9 MB on-device speech-to-text)","org":"Cactus Compute","family":"Needle","released":"2026-10-02","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"Cactus-Compute/whistle","url":"https://huggingface.co/Cactus-Compute/whistle"},{"provider":"GitHub (engine)","url":"https://github.com/Cactus-Compute/needle3"},{"provider":"Browser demo","url":"https://cactuscompute.com/blog/whistle"}],"capabilities":[{"name":"Speech recognition in a 16.9 MB file, on CPU","detail":"One 16.9 MB file with no dependencies, running in the same C++ engine as Cactus's Needle model; ~11 ms to first token; transcribes up to 30 s per pass in English, German, French, Spanish, Italian, Dutch and Polish with language detection, word timestamps and speech embeddings. Vendor benchmarks: ahead of Whisper base on most sets at ~9× smaller and ~6× faster.","first":false,"discovered":"launch","source":"https://cactuscompute.com/blog/whistle"}],"entry":null,"notes":"Aimed at phones, wearables, robots, smart home, cars and microcontrollers. Compared with Whisper base (a small 2022 model), not with large ASR models; vendor-run numbers. Launch post ~167k views; Nobara/GE-Proton maintainer GloriousEggroll reported it made his home assistant respond 'SO damn fast'.","verified":"2026-10-03","body_md":"Launch: https://x.com/cactuscompute/status/2106083041265562075 · Related: [phonon-2](phonon-2.md), [whisper-large-v3](whisper-large-v3.md).","page_url":"https://postcutoff.com/m/cactus-whistle/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":96,"you_url":null,"briefings":null},{"id":"cloudflare-clef","name":"Clef","org":"Cloudflare","family":"Clef","released":"2026-10-01","status":"current","type":"llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":65536,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.24,"unit":"per 1M input tokens (Workers AI); output price not stated in the docs page checked","source":"https://developers.cloudflare.com/workers-ai/models/clef/"},"price_line":"$0.24 in per 1M tokens; output price not published","access":[{"provider":"Cloudflare Workers AI","model_id":"@cf/cloudflare/clef","docs":"https://developers.cloudflare.com/workers-ai/models/clef/"},{"provider":"Hugging Face","url":"https://huggingface.co/Cloudflare/clef"}],"capabilities":[{"name":"Multimodal decision model (typed option scoring)","detail":"Takes text, JSON, images or video plus a schema of typed questions and returns a probability for every allowed option in one prefill-only forward pass; Jev-API compatible; Cloudflare reports 209 ms median latency vs Jev's 524 ms on its 43-benchmark set.","first":false,"discovered":"launch","source":"https://blog.cloudflare.com/clef-decision-models/"}],"entry":"2026-10-01-cloudflare-clef-decision-models","notes":"27B decision model on a Qwen3.8-27B backbone with a vision encoder. Open alternative to TypeSafe's hosted Jev; self-hosting needs ~85 GB VRAM at 64k context (The Register). Benchmarks are Cloudflare's own.","verified":"2026-10-02","body_md":"Open-weights \"decision model\" from Cloudflare: it scores the supplied options instead of generating free text. Use it on Workers AI as `@cf/cloudflare/clef` or download the weights from Hugging Face (Apache 2.0). A smaller, faster sibling is `cloudflare-clef-flash`.","page_url":"https://postcutoff.com/m/cloudflare-clef/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":"https://postcutoff.com/you/cloudflare-clef/","briefings":null},{"id":"cloudflare-clef-flash","name":"Clef-flash","org":"Cloudflare","family":"Clef","released":"2026-10-01","status":"current","type":"llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Cloudflare Workers AI","model_id":"@cf/cloudflare/clef-flash","docs":"https://developers.cloudflare.com/workers-ai/models/clef/"},{"provider":"Hugging Face","url":"https://huggingface.co/Cloudflare/clef-flash"}],"capabilities":[{"name":"Low-latency decision model","detail":"Qwen3.5-9B-based option scorer; Cloudflare reports 38.8 ms median and 122.4 ms p95 latency on its 43-benchmark set and top scores on BFCL (98.76) and API-Bank (93.11).","first":false,"discovered":"launch","source":"https://blog.cloudflare.com/clef-decision-models/"}],"entry":"2026-10-01-cloudflare-clef-decision-models","notes":"Smaller, faster sibling of Clef (Qwen3.5-9B backbone). Self-hosting needs ~41 GB VRAM at 64k context (The Register). Workers AI price not verified.","verified":"2026-10-02","body_md":"Fast open-weights decision model from Cloudflare. Same Jev-compatible API as `cloudflare-clef`. Hugging Face repo name taken from Cloudflare's blog.","page_url":"https://postcutoff.com/m/cloudflare-clef-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":"https://postcutoff.com/you/cloudflare-clef-flash/","briefings":null},{"id":"mai-transcribe-2-streaming","name":"MAI-Transcribe-2-Streaming","org":"Microsoft","family":"MAI-Transcribe","released":"2026-10-01","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.54,"unit":"USD per hour of audio (through end of 2026)","source":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},"price_line":"$0.54 per hour","access":[{"provider":"Microsoft Foundry","docs":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},{"provider":"Azure Voice Live","docs":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},{"provider":"OpenRouter","url":"https://openrouter.ai/"},{"provider":"Web app (MAI Playground)","url":"https://playground.microsoft.ai/"}],"capabilities":[{"name":"Low-latency streaming transcription","detail":"First hypotheses in just over 100 ms; Microsoft claims #1 accuracy on Artificial Analysis for final and partial transcripts, 60 languages with auto-detection.","first":false,"discovered":"launch","source":"https://microsoft.ai/news/our-first-streaming-transcription-model/"}],"entry":"2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1","notes":"Microsoft's first streaming ASR model. Exact API model ids not verified. Benchmarks are Microsoft-reported.","verified":"2026-10-02","body_md":"Source: https://microsoft.ai/news/our-first-streaming-transcription-model/","page_url":"https://postcutoff.com/m/mai-transcribe-2-streaming/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":null,"briefings":null},{"id":"mai-voice-2-1","name":"MAI-Voice-2.1 / MAI-Voice-2.1-Flash","org":"Microsoft","family":"MAI-Voice","released":"2026-10-01","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":22,"unit":"USD per 1M characters (MAI-Voice-2.1); MAI-Voice-2.1-Flash is $15 per 1M characters","source":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},"price_line":"$22 per 1M characters","access":[{"provider":"Microsoft Foundry","docs":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},{"provider":"OpenRouter","url":"https://openrouter.ai/"},{"provider":"Web app (MAI Playground)","url":"https://playground.microsoft.ai/"}],"capabilities":[{"name":"Cross-lingual voice with native accent","detail":"23 languages / 26 locales; a single voice keeps its native accent across languages; consent-gated voice cloning.","first":false,"discovered":"launch","source":"https://microsoft.ai/news/our-first-streaming-transcription-model/"},{"name":"Low-latency Flash variant","detail":"MAI-Voice-2.1-Flash: 150 ms end-to-end latency, 55% faster inference, $15 per 1M characters.","first":false,"discovered":"launch","source":"https://microsoft.ai/news/our-first-streaming-transcription-model/"}],"entry":"2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1","notes":"Successor to MAI-Voice-2 (Build, June 2026). Exact API model ids not verified.","verified":"2026-10-02","body_md":"Source: https://microsoft.ai/news/our-first-streaming-transcription-model/","page_url":"https://postcutoff.com/m/mai-voice-2-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":null,"briefings":null},{"id":"pplx-decider-v1-27b","name":"pplx-decider-v1-27b","org":"Perplexity","family":"pplx-decider","released":"2026-10-01","status":"current","type":"llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.04,"output":0,"unit":"per 1M tokens (output tokens free; Decisions API)","source":"https://docs.perplexity.ai/docs/decisions/quickstart"},"price_line":"$0.04 in per 1M tokens; output free","access":[{"provider":"Perplexity API","model_id":"pplx-decider-v1-27b","endpoint":"https://api.perplexity.ai/v1/decisions","docs":"https://docs.perplexity.ai/docs/decisions/quickstart"},{"provider":"Hugging Face","url":"https://huggingface.co/perplexity-ai/pplx-decider-v1-27b"}],"capabilities":[{"name":"Decision model (probabilities over fixed answers)","detail":"Returns a probability distribution over the allowed options for 1–128 typed questions per request instead of generating text; Perplexity reports 85.71% on its 11-benchmark panel vs 84.51% for Jev.","first":false,"discovered":"launch","source":"https://huggingface.co/perplexity-ai/pplx-decider-v1-27b"}],"entry":"2026-10-01-perplexity-decisions-api-pplx-decider","notes":"Fine-tuned from Qwen3.8-27B. Context is the API's 262,144-token input limit. Benchmarks are Perplexity's own; Jev still leads on 4–6 of 11 tests.","verified":"2026-10-03","body_md":"Open-weights decision model behind Perplexity's Decisions API. Send a request with your content plus typed questions to `POST https://api.perplexity.ai/v1/decisions`; output tokens are free.","page_url":"https://postcutoff.com/m/pplx-decider-v1-27b/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":"https://postcutoff.com/you/pplx-decider-v1-27b/","briefings":null},{"id":"strands-decider-2b","name":"Strands Decider 2B","org":"Amazon Web Services (Strands Labs)","family":"Strands Decider","released":"2026-10-01","status":"preview","type":"llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0,"output":0,"unit":"free open weights (self-hosted)","source":"https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19"},"price_line":"Free","access":[{"provider":"Hugging Face","url":"https://huggingface.co/StrandsAgents/strands-decider-2B-hobson-v19"}],"capabilities":[{"name":"Decision model (option scoring)","detail":"Replaces the LM head of Qwen3.5-2B-Base with a readout/pointer head that scores the supplied options and returns calibrated confidences in one pass: JevBench public 0.723, ~106 ms median latency on an RTX 3090 (VentureBeat).","first":false,"discovered":"launch","source":"https://venturebeat.com/technology/amazon-unveils-a-free-fast-open-source-jev-killer-strands-decider-2b-makes-decisions-in-fractions-of-a-second"}],"entry":"2026-09-15-typesafe-jev-system-one-model","notes":"Experimental release from AWS Strands Labs (LoRA adapter on Qwen/Qwen3.5-2B-Base). Uses: routing, tool selection, guardrails, triage. Weaker on long multi-step documents (model card). Open-source alternative to TypeSafe's hosted Jev.","verified":"2026-10-01","body_md":"Open-weights \"decision model\" from AWS Strands Labs. It picks among given options or rates on a scale with calibrated probabilities; it does not generate free text. Download from Hugging Face (Apache 2.0).","page_url":"https://postcutoff.com/m/strands-decider-2b/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":"https://postcutoff.com/you/strands-decider-2b/","briefings":null},{"id":"tavus-griffin","name":"Tavus Griffin (Human Interaction Model)","org":"Tavus","family":"Griffin","released":"2026-10-01","status":"preview","type":"video-gen","modality_in":["video","audio"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Tavus (research preview)","url":"https://www.tavus.io/griffin"}],"capabilities":[{"name":"Full-duplex face-to-face video conversation","detail":"Video-to-video model that sees and hears the user while speaking, interrupts and takes interruptions, and reacts to gestures and objects in real time; Tavus calls it the first 'Human Interaction Model'.","first":true,"discovered":"launch","source":"https://www.tavus.io/griffin"},{"name":"'Video Turing test' result (vendor study)","detail":"48% of 54 participants in 1-minute blind live calls judged it human (Tavus's earlier stack ≤2%); NVIDIA VideoFDB generation 3.83/5 vs human 3.92; median latency ~1.9 s.","first":false,"discovered":"launch","source":"https://www.tavus.io/griffin"}],"entry":"2026-10-01-tavus-griffin-video-turing-test","notes":"Research preview; no public API id or pricing found. Benchmarks: Tavus's own study plus NVIDIA's VideoFDB leaderboard (Griffin-Lite).","verified":"2026-10-02","body_md":null,"page_url":"https://postcutoff.com/m/tavus-griffin/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":132,"you_url":null,"briefings":null},{"id":"cohere-embed-v5-fast","name":"Cohere Embed 5 Fast","org":"Cohere","family":"Embed","released":"2026-09-30","status":"current","type":"embedding","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.08,"unit":"per 1M text tokens (images $0.40 per 1M tokens)","source":"https://cohere.com/blog/embed-5"},"price_line":"$0.08 in per 1M tokens","access":[{"provider":"Cohere API","model_id":"embed-v5.0-fast","endpoint":"https://api.cohere.com/v2/embed","docs":"https://docs.cohere.com/changelog/embed-v5"},{"provider":"Microsoft Foundry"},{"provider":"Amazon SageMaker"},{"provider":"Cohere Model Vault"}],"capabilities":[{"name":"Shared embedding space across tiers","detail":"Embed 5 Pro and Fast vectors are interoperable: index with Pro, query with Fast without re-indexing.","first":false,"discovered":"launch","source":"https://docs.cohere.com/changelog/embed-v5"},{"name":"Multimodal 128K embeddings","detail":"Text, image and fused text+image input, 100+ languages, Matryoshka dims 256-2048.","first":false,"discovered":"launch","source":"https://docs.cohere.com/changelog/embed-v5"}],"entry":"2026-09-30-cohere-embed-5","notes":"Latency/cost tier for live queries; ~2.4x Pro throughput; ViDoRe V3 84.5 (Cohere). Output is vectors (modality_out text used as placeholder). Bedrock/Foundry model ids not verified.","verified":"2026-10-01","body_md":"```bash\ncurl https://api.cohere.com/v2/embed -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"embed-v5.0-fast\",\"input_type\":\"search_query\",\"embedding_types\":[\"float\"],\"texts\":[\"hello world\"]}'\n```","page_url":"https://postcutoff.com/m/cohere-embed-v5-fast/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":165,"you_url":null,"briefings":null},{"id":"cohere-embed-v5-pro","name":"Cohere Embed 5 Pro","org":"Cohere","family":"Embed","released":"2026-09-30","status":"current","type":"embedding","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.12,"unit":"per 1M text tokens (images $0.40 per 1M tokens)","source":"https://cohere.com/blog/embed-5"},"price_line":"$0.12 in per 1M tokens","access":[{"provider":"Cohere API","model_id":"embed-v5.0-pro","endpoint":"https://api.cohere.com/v2/embed","docs":"https://docs.cohere.com/changelog/embed-v5"},{"provider":"Microsoft Foundry"},{"provider":"Amazon SageMaker"},{"provider":"Cohere Model Vault"}],"capabilities":[{"name":"Shared embedding space across tiers","detail":"Embed 5 Pro and Fast vectors are interoperable: index with Pro, query with Fast without re-indexing.","first":false,"discovered":"launch","source":"https://docs.cohere.com/changelog/embed-v5"},{"name":"Multimodal 128K embeddings","detail":"Text, image and fused text+image input, 100+ languages, Matryoshka dims 256-2048.","first":false,"discovered":"launch","source":"https://docs.cohere.com/changelog/embed-v5"}],"entry":"2026-09-30-cohere-embed-5","notes":"Highest-quality tier for offline indexing; ViDoRe V3 85.8 (Cohere). Output is vectors (modality_out text used as placeholder). Bedrock/Foundry model ids not verified.","verified":"2026-10-01","body_md":"```bash\ncurl https://api.cohere.com/v2/embed -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"embed-v5.0-pro\",\"input_type\":\"search_query\",\"embedding_types\":[\"float\"],\"texts\":[\"hello world\"]}'\n```","page_url":"https://postcutoff.com/m/cohere-embed-v5-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":165,"you_url":null,"briefings":null},{"id":"gemini-4-argon","name":"Gemini 4 Argon","org":"Google DeepMind","family":"Gemini 4","released":"2026-09-30","status":"preview","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":1000000,"knowledge_cutoff":null,"pricing":{"input":2,"output":10,"cached_input":0.1,"unit":"per 1M tokens (introductory 50% discount, end date not announced; standard price $4 input / $20 output; cached input 95% off input)","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"},"price_line":"$2 in, $10 out per 1M tokens","access":[{"provider":"Google DeepMind Fairwind Program (vetted cyber defenders only; standalone or with the CodeMender agent)","url":"https://deepmind.google/fairwind-program/"},{"provider":"Fairwind access form","url":"https://rsvp.withgoogle.com/events/fairwind-program-interest-form"},{"provider":"Gemini API / Google AI Ultra (announced as next, no date)","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"}],"capabilities":[{"name":"1M-token output limit","detail":"Can generate up to 1M output tokens in one response or trajectory (previous Gemini models: 64K). The Gemini API's new Long Decode Continuation feature pauses and resumes long responses across calls to avoid timeouts (per Artificial Analysis).","first":true,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"},{"name":"Enterprise knowledge work","detail":"Vendor-reported: Vals Index 68.9% (#1; confirmed on the Vals AI leaderboard), Harvey Legal Agent Benchmark 19.6% vs 6.7% for the next model, Vals Finance Agent v2 65.4%, Zapier AutomationBench 51.3% (#1).","first":false,"discovered":"launch","source":"https://deepmind.google/models/gemini/"},{"name":"Low hallucination rate","detail":"Artificial Analysis measured a 15% hallucination rate on AA-Omniscience, the lowest of any model scoring 45+ on its Intelligence Index (GPT-6 Astra 51%).","first":false,"discovered":"launch","source":"https://x.com/ArtificialAnlys/status/2105392625788637299"},{"name":"Frontier software engineering (mixed)","detail":"Vendor-reported DeepSWE v1.1 77.9% (Opus 5.5 74.2%, GPT-6 Astra 74.1%), but it trails on FrontierSWE v2 (55.0% vs Astra 65.5%) and Terminal-Bench 4.0 (57.4% vs Opus 5.5 66.4%). Migrating C/C++ to Rust across Google, up to 800K+ lines.","first":false,"discovered":"launch","source":"https://deepmind.google/models/gemini/"},{"name":"Autonomous vulnerability discovery and patching","detail":"CWE-bench v1 68% pass@1, tied first with GPT-6 Astra and Grok 4.7 (public leaderboard). Offered without cyber guardrails to Fairwind defenders. Google reports 85.8% on its internal vulnerability benchmark and 70.9% on Wiz's pentest benchmark.","first":false,"discovered":"launch","source":"https://cwe-bench.com/"},{"name":"Long-context and long-video understanding","detail":"Vendor-reported GraphWalks BFS 256K–1M 84.2% (Astra 71.8%) and LVBench 91.7% (state of the art at launch).","first":false,"discovered":"launch","source":"https://deepmind.google/models/gemini/"}],"entry":"2026-09-30-gemini-4-argon","notes":"Announced 2026-09-30 with staged access: Fairwind Program first, then paid API customers and Google AI Ultra 'as soon as possible'. Independent launch-day results: Artificial Analysis Intelligence Index 53 (tied with GPT-6 Astra), Arena Text #1 (1525). No public API model id, knowledge cutoff or model card as of 2026-09-30. Arena lists it as 'gemini-4-argon-high'. The 1M context comes from Artificial Analysis and Arena, not from Google. modality_in follows the Gemini family: AA's model page lists text + image, and its X post says text, image, video and speech. 'Argon' replaces the '3.x Pro' naming. Zero data retention is available for Fairwind partners using it as a managed model.","verified":"2026-10-01","body_md":"Google's first Gemini 4 model. It is not generally available yet, and there is no Gemini API or Vertex AI model id in the docs\n(checked 2026-09-30). Update this file with the model id, knowledge cutoff, docs link and model card once it reaches the Gemini API.\n\nBenchmark highlights (Google's table at deepmind.google/models/gemini, with methodology in\n[this PDF](https://storage.googleapis.com/deepmind-media/gemini/gemini_4_argon_model_evaluation.pdf)). Scores are for the highest\nthinking setting, and rival scores are mostly the providers' own figures.\n\n| Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |\n|---|---|---|---|---|\n| Vals Index | 68.9 | 63.1 | 65.8 | 67.0 |\n| DeepSWE v1.1 | 77.9 | 74.1 | 67.4 | 74.2 |\n| FrontierSWE v2 | 55.0 | 65.5 | 56.3 | 62.3 |\n| Terminal-Bench 4.0 | 57.4 | 58.2 | 57.9 | 66.4 |\n| Terminal-Bench Science 0.1 | 57.6 | 68.1 | 52.6 | 63.3 |\n| GraphWalks 256K–1M | 84.2 | 71.8 | 65.0 | 66.8 |\n| OSWorld-2.0 (offline, partial) | 69.2 | 72.6 | — | — |\n| LVBench | 91.7 | 87.5 | 79.7 | 83.7 |\n| CWE-bench v1 | 68.0 | 68.0 | 58.0 | 67.0 |","page_url":"https://postcutoff.com/m/gemini-4-argon/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":165,"you_url":"https://postcutoff.com/you/gemini-4-argon/","briefings":null},{"id":"praxis-1","name":"Praxis-1","org":"Runway","family":"Praxis","released":"2026-09-30","status":"preview","type":"robotics","modality_in":["video","image","text"],"modality_out":["action"],"open_weights":false,"model_license":"not yet stated (open weights promised 'in the coming months')","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Runway early access (request form)","url":"https://runway.com/research/introducing-praxis-1"}],"capabilities":[{"name":"World action model pretrained on web video","detail":"A generalist robot policy built on Runway's large-scale video pretraining; Runway reports that after fine-tuning, policies pretrained on web video and on teleoperated robot video reached nearly the same placement error (16.1 cm vs 16.0 cm over 93 evaluation pairs).","first":false,"discovered":"launch","source":"https://runway.com/research/introducing-praxis-1"},{"name":"Cross-embodiment policy","detail":"One policy for bimanual rigs, 6-DoF arms and mobile bases without retraining; early partners Noble Machines, Standard Bots and Ultra run it on their own hardware.","first":false,"discovered":"launch","source":"https://www.therobotreport.com/runway-introduces-praxis-1-world-action-model-robotics/"},{"name":"Policy evaluation inside a world model","detail":"Runway says simulating robot policies inside its world model predicted real-world results with 0.95 correlation.","first":false,"discovered":"launch","source":"https://runway.com/research/introducing-praxis-1"}],"entry":"2026-09-30-runway-praxis-1","notes":"Early access only as of 2026-10-03; Runway says the weights will be released openly 'in the coming months' (tracked in data/upcoming/2026-10-01-runway-praxis-1-open-weights.md). Parameter count, license and Hugging Face id not yet published. Runway names transparent materials and deformable cloth as weak spots.","verified":"2026-10-03","body_md":"Runway's first robotics model, announced around 2026-09-30 by Head of Robotics Andy Chen. Set `open_weights: true` and add the Hugging Face id when the weights ship.\n\nSources: [Runway Research: Introducing Praxis-1](https://runway.com/research/introducing-praxis-1), [The Robot Report](https://www.therobotreport.com/runway-introduces-praxis-1-world-action-model-robotics/), [Runway on X](https://x.com/runwayml/status/2105419393928929671).","page_url":"https://postcutoff.com/m/praxis-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":165,"you_url":null,"briefings":null},{"id":"assemblyai-universal-3-6-pro-realtime","name":"AssemblyAI Universal-3.6 Pro Realtime","org":"AssemblyAI","family":"Universal-3","released":"2026-09-29","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.45,"unit":"per hour of streaming session (same as Universal-3.5 Pro Realtime); Universal-Streaming EN/multi $0.15/hr","source":"https://www.assemblyai.com/pricing"},"price_line":"$0.45 per hour","access":[{"provider":"AssemblyAI API","model_id":"universal-3-6-pro","endpoint":"wss://streaming.assemblyai.com/v3/ws?model=universal-3-6-pro","docs":"https://www.assemblyai.com/blog/universal-3-6-pro-realtime"},{"provider":"AssemblyAI Voice Agent API","model_id":"(default STT)","endpoint":"wss://agents.assemblyai.com/v1/ws"}],"capabilities":[{"name":"Promptable streaming STT for voice agents","detail":"Prompting + keyterms together, real-time diarization, entity-aware endpointing and native code-switching in 32 languages with auto language detection; 5.13% normalized WER (vs 5.80% for 3.5 Pro Realtime), short-response WER 1.45%; median endpoint latency 537 ms.","first":false,"discovered":"launch","source":"https://www.assemblyai.com/blog/universal-3-6-pro-realtime"}],"entry":null,"notes":"Lineage: Universal-3 Pro Streaming (Mar 2026) -> Universal-3.5 Pro Realtime (2026-06-23) -> 3.6 (2026-09-29). Older streaming ids u3-rt-pro/u3-pro replaced. Voice Agent API ($4.50/hr all-in: STT+LLM+TTS) GA April 2026. AssemblyAI roadmap targets 30+ native languages for the next Universal-3.x in Q4 2026.","verified":"2026-09-29","body_md":"Sources: https://www.assemblyai.com/blog/universal-3-6-pro-realtime , https://www.assemblyai.com/llms/models.md , https://www.assemblyai.com/pricing","page_url":"https://postcutoff.com/m/assemblyai-universal-3-6-pro-realtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":213,"you_url":null,"briefings":null},{"id":"gpt-6-1-sol","name":"GPT-6.1 Sol","org":"OpenAI","family":"GPT-6","released":"2026-09-29","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-04","pricing":{"input":2,"cached_input":0.1,"output":10,"cache_write":2.5,"unit_note":"prompts >272K input tokens billed at 2x input/cache and 1.5x output; Fast mode 2x; Batch/Flex 50% off; Ultrafast not yet priced","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/models/gpt-6.1-sol"},"price_line":"$2 in, $10 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-6.1-sol","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-6.1-sol"},{"provider":"OpenRouter","model_id":"openai/gpt-6.1-sol","url":"https://openrouter.ai/openai/gpt-6.1-sol"},{"provider":"GitHub Copilot","url":"https://github.blog/changelog/2026-09-29-gpt-6-1-sol-in-github-copilot"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Near-Astra quality at one-fifth the price","detail":"OpenAI says it nearly matches GPT-6 Astra on agentic coding (DeepSWE v1.1), computer use (OSWorld 2.0, within 2.1 pts at ~1/7 cost/task) and professional work at $2/$10 vs Astra's $10/$50.","first":false,"discovered":"launch","source":"https://openai.com/index/introducing-gpt-6-1-sol/"},{"name":"Near-Astra ARC-AGI-3 score at lower cost","detail":"ARC Prize verified (Sept 30, 2026): ARC-AGI-3 96.4% at $4.4K with the provider-adapter harness (52.7% with the standard harness) vs GPT-6 Astra's 99.9%, at 77% lower cost; ARC-AGI-2 94.2% at $0.25/task.","first":false,"discovered":"later","source":"https://x.com/arcprize/status/2105391846067249649"},{"name":"95% cached-input discount","detail":"Cached input costs $0.10 per 1M tokens (5% of the input rate), half of GPT-6 Sol's cached price.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6.1-sol"},{"name":"Improved alignment vs GPT-6 Sol","detail":"Failed to disclose a broken search tool in 2.1% of adversarial tests (GPT-6 Sol 4.9%, Astra 1.5%); no observed attempts to bypass the automated safety reviewer.","first":false,"discovered":"launch","source":"https://cdn.openai.com/pdf/38e3efcf-545e-44cd-99ec-2b7eb395f4cc/oai_GPT_6_1_Sol.pdf"},{"name":"Full hosted tool suite","detail":"Responses API tools: web/file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, tool search.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6.1-sol"},{"name":"Multi-agent in the Responses API (beta)","detail":"The model can spin up and coordinate parallel subagents inside one Responses API request and synthesize their results; enable with the beta header OpenAI-Beta: responses_multi_agent=v1 (also available for all GPT-5.6 models). OpenAI's Oct 2 2026 GPT-6 model guide names GPT-6.1 Sol as the family member that supports delegation.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/guides/responses-multi-agent"}],"entry":"2026-09-29-gpt-6-1-sol","notes":"Successor to GPT-6 Sol, released one week later at DevDay 2026. reasoning.effort low/medium(default)/high/xhigh/max (no none/minimal). Max input 922K tokens. Chat Completions supported without tool calling. US/EU data residency (Fast mode unavailable with EU residency). In ChatGPT Work and Codex for Plus and above; not yet in Chat as of launch. Ultrafast version promised 'in the coming days'; as of Oct 2 2026 OpenAI's GPT-6 model guide and the Ultrafast docs still list Ultrafast for GPT-6 Astra only. OpenRouter also lists openai/gpt-6.1-sol-pro.","verified":"2026-10-02","body_md":"OpenAI's default mid-tier model for coding agents, computer use and document work as of Sept 29, 2026.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-6.1-sol\",\"reasoning\":{\"effort\":\"medium\"},\"input\":\"Hello\"}'\n```\n\nRate limits (standard): Tier 1 500 RPM / 500K TPM … Tier 5 15,000 RPM / 40M TPM.\n\nSources:\n- https://openai.com/index/introducing-gpt-6-1-sol/\n- https://developers.openai.com/api/docs/models/gpt-6.1-sol\n- https://developers.openai.com/api/docs/pricing\n- https://github.blog/changelog/2026-09-29-gpt-6-1-sol-in-github-copilot\n- https://openai.com/index/practical-guide-building-gpt-6 (GPT-6 family model guide, Oct 2 2026)\n- https://developers.openai.com/api/docs/guides/responses-multi-agent","page_url":"https://postcutoff.com/m/gpt-6-1-sol/","events_after":725,"major_after":180,"historic_after":38,"missing_at_launch":478,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-6-1-sol/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-04-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-04-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-04-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-04.md"}},{"id":"ling-3-1-flash","name":"Ling 3.1 Flash","org":"inclusionAI (Ant Group)","family":"Ling 3","released":"2026-09-29","status":"preview","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":262144,"max_output":32768,"knowledge_cutoff":null,"pricing":{"input":0,"output":0,"unit":"per 1M tokens (USD); free trial on gateways (Vercel lists free through 2026-10-13)","source":"https://openrouter.ai/inclusionai/ling-3.1-flash"},"price_line":"Free","access":[{"provider":"OpenRouter","model_id":"inclusionai/ling-3.1-flash","url":"https://openrouter.ai/inclusionai/ling-3.1-flash"},{"provider":"Vercel AI Gateway","model_id":"inclusionai/ling-3.1-flash","url":"https://vercel.com/ai-gateway/models/ling-3.1-flash"}],"capabilities":[{"name":"560B-total / 25B-active hybrid-reasoning MoE","detail":"Sparse MoE for coding, agents and tool use; reasoning enabled by default (can be turned off via the reasoning parameter on OpenRouter); tools and tool_choice supported.","first":false,"discovered":"launch","source":"https://openrouter.ai/inclusionai/ling-3.1-flash"}],"entry":"2026-09-29-ant-inclusionai-ling-3-1-flash","notes":"Trial capped at 256K context; inclusionAI says it will enable ~1M context and open-source the model after the two-week trial (press, not an official page). No weights, license or official pricing yet; update open_weights/license/pricing when released. Served on Vercel by Novita.","verified":"2026-10-02","body_md":"Free trial model as of Oct 2, 2026. Example (OpenRouter, OpenAI-compatible):\n\n```bash\ncurl https://openrouter.ai/api/v1/chat/completions -H \"Authorization: Bearer $OPENROUTER_API_KEY\" \\\n -H \"Content-Type: application/json\" -d '{\"model\":\"inclusionai/ling-3.1-flash\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```","page_url":"https://postcutoff.com/m/ling-3-1-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":213,"you_url":"https://postcutoff.com/you/ling-3-1-flash/","briefings":null},{"id":"claude-sonnet-5-5","name":"Claude Sonnet 5.5","org":"Anthropic","family":"Claude 5","released":"2026-09-28","status":"current","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-06","pricing":{"input":2,"output":10,"cache_read":0.2,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$2 in, $10 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-sonnet-5-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/sonnet-5-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-sonnet-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-sonnet-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-sonnet-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-sonnet-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-sonnet-5.5","url":"https://openrouter.ai/anthropic/claude-sonnet-5.5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Opus-level knowledge work at Sonnet price","detail":"Scores nearly level with Opus 5.5 on GDPval-AA (1844 vs 1846 Elo), at $2/$10 per MTok.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-sonnet-5-5"},{"name":"Large agentic-coding jump","detail":"Anthropic reports 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5, and up to 30% lower cost per task.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-sonnet-5-5"},{"name":"Beat Pokemon Red from screenshots","detail":"Anthropic says it is the first Sonnet model to finish Pokemon Red using only screenshots.","first":true,"discovered":"launch","source":"https://www.anthropic.com/claude-sonnet-5-5"},{"name":"between_tools thinking mode","detail":"New thinking type that turns off up-front thinking while still reasoning between tool calls; it replaces thinking: disabled.","first":true,"discovered":"launch","source":"https://platform.claude.com/docs/en/models/sonnet-5-5/overview"},{"name":"Token efficiency","detail":"A Balyasny test used 121k tokens per task, versus 497k for Sonnet 5.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-sonnet-5-5"}],"entry":null,"notes":"Best speed/intelligence balance. Adaptive thinking on by default (effort default high); thinking {type: disabled} returns 400, use {type: between_tools} at effort high or below; forced tool_choice any/tool returns 400; non-default temperature/top_p/top_k return 400. Batch $1/$5.","verified":"2026-09-29","body_md":"Everyday coding, agents and enterprise workloads at Sonnet pricing; successor to Claude Sonnet 5 at the same price.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-sonnet-5-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/sonnet-5-5/overview\n- Announcement: https://www.anthropic.com/claude-sonnet-5-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-sonnet-5-5/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":388,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-sonnet-5-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"elevenlabs-v4","name":"Eleven v4 / Eleven v4 Turbo","org":"ElevenLabs","family":"Eleven v4","released":"2026-09-28","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.08,"unit":"per 1K characters, list price for eleven_v4 (eleven_v4_turbo list $0.04/1K). Launch promo 72% off until 2026-10-12: eleven_v4 $0.022/1K, eleven_v4_turbo $0.011/1K (= $22 / $11 per 1M chars)","promo_per_1k_characters":0.022,"turbo_per_1k_characters":0.04,"turbo_promo_per_1k_characters":0.011,"promo_ends":"2026-10-12","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.08 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_v4","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/overview/capabilities/text-to-speech/eleven-v4"},{"provider":"ElevenLabs API","model_id":"eleven_v4_turbo","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/models"},{"provider":"fal","model_id":"elevenlabs/tts/eleven-v4","docs":"https://fal.ai/models/elevenlabs/tts/eleven-v4"},{"provider":"fal","model_id":"elevenlabs/tts/eleven-v4-turbo","docs":"https://fal.ai/models/elevenlabs/tts/eleven-v4-turbo"},{"provider":"Web app (ElevenCreative)","url":"https://elevenlabs.io/app"},{"provider":"Landing page / demos","url":"https://elevenlabs.io/v4"}],"capabilities":[{"name":"Context-aware \"performed\" delivery (new architecture)","detail":"Entirely new TTS architecture that 'reads a script the way a voice actor would', interpreting tone, pacing, emotion, character and context; preferred by ~75% of listeners (65-81% range) in blind head-to-head tests vs Cartesia Sonic 3.6, Inworld TTS-2, Gemini TTS, xAI TTS and GPT-4o mini TTS.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/eleven-v4"},{"name":"#1 on Artificial Analysis TTS arena","detail":"Took #1 on the Artificial Analysis Provider Voice TTS Arena (Elo ~1315-1319 at launch, ahead of Cartesia Sonic 3.6 at 1275 and Gemini 3.8 Flash TTS at 1267) and #1 on AA's Pronunciation Robustness benchmark, #2 on Controlled Voice.","first":false,"discovered":"launch","source":"https://artificialanalysis.ai/text-to-speech/leaderboard"},{"name":"Real-time Turbo variant (~100 ms)","detail":"eleven_v4_turbo: ~100 ms median inference latency, ~150 ms median time to first speech (vs Cartesia Sonic 3.6 262 ms, GPT-4o mini TTS 814 ms per ElevenLabs), for voice agents.","first":false,"discovered":"launch","source":"https://elevenlabs.io/v4"},{"name":"Cross-lingual native accent, 90+ languages","detail":"90+ languages (new: Cantonese, Mongolian, Odia); when target language differs from the reference voice, v4 speaks with a fluent native accent instead of carrying over the source accent.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/overview/capabilities/text-to-speech/eleven-v4"},{"name":"Inline tags incl. sound effects and free-text direction","detail":"Inline tags direct delivery, emotion, pacing, reactions, SFX and style, e.g. [laughs], [said angrily in French accent], [light rain], [phone buzzing], [quick, light, playful pace].","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/eleven-v4"},{"name":"Voice cloning from 10 s, PVC support restored","detail":"Instant Voice Clones from ~10 s of audio (docs still recommend 1-2 min); Professional Voice Clones supported again (not available on v3); speaker identity kept across regenerations/long-form.","first":false,"discovered":"launch","source":"https://elevenlabs.io/v4"},{"name":"IPA pronunciation control","detail":"Pronunciation control with IPA support; more natural multi-speaker dialogue.","first":false,"discovered":"launch","source":"https://www.youtube.com/watch?v=th_tXR2QQ6U"}],"entry":"2026-09-28-elevenlabs-eleven-v4","notes":"Launched 2026-09-28 (blog, YouTube 07:01 PT, X) in ElevenAgents, ElevenCreative and ElevenAPI, incl. free tier. eleven_v4: 10,000 chars/request; eleven_v4_turbo: no char limit listed on models page. Output formats MP3, WAV/PCM, u-law. Limitations: no Style/Speed sliders, no SSML (Stability + Similarity only); Voice Design voices may perform worse than with earlier models. Launch promo also: v4 free for Creator+ plans in ElevenCreative up to 2x monthly credits for two weeks. Third-party: on fal since launch day (fal X post https://x.com/fal/status/2104630460542325071): elevenlabs/tts/eleven-v4 at $0.08/1K chars and elevenlabs/tts/eleven-v4-turbo at $0.04/1K chars, the ElevenLabs list prices (fal model pages, checked 2026-09-29). Research led by Piotr Dabkowski (per press).","verified":"2026-09-29","body_md":"ElevenLabs' newest, most expressive TTS (eleven_v4); eleven_v4_turbo for agents/real time. Successor to Eleven v3.\n\n```bash\ncurl -X POST \"https://api.elevenlabs.io/v1/text-to-speech/$VOICE_ID\" -H \"xi-api-key: $ELEVENLABS_API_KEY\" \\\n  -H \"Content-Type: application/json\" -d '{\"text\":\"[laughs] Well... I did not expect that!\",\"model_id\":\"eleven_v4\"}' -o out.mp3\n```\n\nLaunch videos: https://www.youtube.com/watch?v=th_tXR2QQ6U (ElevenLabs), https://www.youtube.com/watch?v=4QHFkK2MTcw (ElevenLabs Developers).\n\nSources: https://elevenlabs.io/blog/eleven-v4 , https://elevenlabs.io/v4 , https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://x.com/ElevenLabs/status/2104572127617994917 , https://x.com/ElevenLabs/status/2104572138347004161 , https://x.com/ArtificialAnlys/status/2104578736687653293\n\n## Changelog\n- 2026-09-29: added fal access rows (elevenlabs/tts/eleven-v4, eleven-v4-turbo) with prices from fal model pages","page_url":"https://postcutoff.com/m/elevenlabs-v4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":247,"you_url":null,"briefings":null},{"id":"phonon-2","name":"Phonon-2 (open weights, on-device ASR)","org":"Fermion Research","family":"Phonon","released":"2026-09-28","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"cc-by-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"FermionResearch/Phonon-2","url":"https://huggingface.co/FermionResearch/Phonon-2"},{"provider":"PyPI (CLI)","model_id":"fermion-research","url":"https://pypi.org/project/fermion-research/","docs":"https://github.com/fermionresearch/phonon"},{"provider":"Docker","model_id":"ghcr.io/fermionresearch/phonon-cpu:2.0.2"},{"provider":"Mac app (Detta)","url":"https://www.fermionresearch.com/products/detta"}],"capabilities":[{"name":"~2-bit quantised Parakeet for on-device English ASR","detail":"A quantisation-aware-trained compression of NVIDIA parakeet-tdt-0.6b-v3: encoder weights at one of five learned levels (~2.1 bits), a 164 MB download vs the 2.5 GB teacher, averaging 5.21% WER on the Open ASR Leaderboard's seven English sets (teacher 4.96%, Whisper large-v3-turbo 6.58%, vendor-run numbers).","first":false,"discovered":"launch","source":"https://huggingface.co/FermionResearch/Phonon-2"},{"name":"Fast local transcription","detail":"Vendor figures: about 174x realtime on an M5 MacBook Air (an hour of audio in ~20 s), 143x on eight Zen 5 cores, 6,680x on one H100 at batch 128.","first":false,"discovered":"launch","source":"https://huggingface.co/FermionResearch/Phonon-2"}],"entry":null,"notes":"English only. Not a new architecture: a compressed NVIDIA Parakeet TDT 0.6B v3 (tokenizer and output conventions unchanged). The launch post's 'more accurate than Whisper large at 1/10 the size' refers to whisper-large-v3-turbo in the vendor's own table; benchmarks not independently reproduced. Fermion Research is a small startup (founder Manan Gupta).","verified":"2026-09-30","body_md":"Launch post: https://x.com/yoitsmanan/status/2104990913886031993 (≈37k views). Related: [nvidia-parakeet-canary](nvidia-parakeet-canary.md), [whisper-large-v3](whisper-large-v3.md).","page_url":"https://postcutoff.com/m/phonon-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":247,"you_url":null,"briefings":null},{"id":"qwen-audio-3-1-asr","name":"Qwen-Audio-3.1-ASR (Flash)","org":"Alibaba (Qwen)","family":"Qwen-Audio 3.1","released":"2026-09-23","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Alibaba Cloud Model Studio (streaming)","model_id":"qwen-audio-3.1-asr-flash-streaming","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"Alibaba Cloud Model Studio / QwenCloud (file transcription)","model_id":"qwen-audio-3.1-asr-flash-filetrans","docs":"https://www.qwencloud.com/models/qwen-audio-3.1-asr-flash-filetrans"}],"capabilities":[{"name":"Multilingual + dialect ASR with disfluency cleanup","detail":"Improved multilingual and Chinese-dialect recognition that automatically removes filler words and repetitions; launched with up to 95% price cut.","first":false,"discovered":"launch","source":"https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/"}],"entry":"2026-09-23-qwen-audio-3-1","notes":"Secondary sources report 30 languages + Chinese dialects and ~160 ms latency (unverified). Sibling Qwen-Audio-3.1-ASR-Next adds speaker diarization with timestamps, emotion and sound-event detection (API id not verified). Previous: qwen-audio-3.0-asr-flash; open-weights alternative Qwen3-ASR (see qwen3-asr). Pricing not verified on an official page.","verified":"2026-09-29","body_md":"Hosted speech recognition in the Qwen-Audio 3.1 stack.\n\nSources: https://www.alibabacloud.com/help/en/model-studio/models · https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/","page_url":"https://postcutoff.com/m/qwen-audio-3-1-asr/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":322,"you_url":null,"briefings":null},{"id":"qwen-audio-3-1-realtime","name":"Qwen-Audio-3.1-Realtime (Plus)","org":"Alibaba (Qwen)","family":"Qwen-Audio 3.1","released":"2026-09-23","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":262144,"max_output":16384,"knowledge_cutoff":null,"pricing":{"audio_input":6.4,"text_input":0.8,"text_output":6.4,"audio_output":24,"unit":"USD per 1M tokens (QwenCloud list price; audio/text output $24 when audio is generated, $6.40 text-only)","source":"https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus"},"price_line":"$0.80 text in, $6.40 text out per 1M tokens","access":[{"provider":"QwenCloud (Realtime WebSocket)","model_id":"qwen-audio-3.1-realtime-plus","endpoint":"wss://maas.qwencloudapi.com/api-ws/v1/realtime?model=qwen-audio-3.1-realtime-plus","docs":"https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus"},{"provider":"Alibaba Cloud Model Studio (Singapore / Beijing)","model_id":"qwen-audio-3.1-realtime-plus","endpoint":"wss://{WorkspaceId}.sg-singapore.maas.aliyuncs.com/api-ws/v1/realtime","docs":"https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides"}],"capabilities":[{"name":"Full-duplex agentic voice (\"Think, Act, Speak and Coordinate\")","detail":"Listens while speaking, decides whether to keep listening, speak, stop or resume; function calling and built-in web search. Task success 82.0% vs 78.4% for the previous version; replies to background speech fell from 73.0% to 13.0% (Full-Duplex-Bench v1.5).","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2609.25176"},{"name":"Three turn-taking modes and voice cloning","detail":"server_vad, semantic smart_turn and push-to-talk modes; system voices plus cloned custom voices; 16 kHz PCM in, 24 kHz PCM out.","first":false,"discovered":"launch","source":"https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides"},{"name":"~85% price cut at launch","detail":"Alibaba cut Realtime prices about 85% with the 3.1 release (TTS ~70%, ASR up to 95%).","first":false,"discovered":"launch","source":"https://x.com/Alibaba_Qwen/status/2102687258990026993"}],"entry":"2026-09-23-qwen-audio-3-1","notes":"Languages: de, en, es, fr, id, it, ja, ko, pt, ru, zh (Mandarin, Cantonese and 18+ Chinese varieties). Predecessors qwen-audio-3.0-realtime-plus / -flash (July 2026) still listed. Press (MarkTechPost) reports interruption-stop latency 1.116 s vs 0.383 s for GPT-Realtime-2 and higher red-team refusal for GPT-Realtime-2; not verified on an official page. Release date is the announcement date (Qwen X post / Apsara); Model Studio pricing for this id not verified.","verified":"2026-09-29","body_md":"Alibaba's hosted real-time voice agent model (WebSocket Realtime API, OpenAI-Realtime-style events).\n\nSources: https://www.qwencloud.com/models/qwen-audio-3.1-realtime-plus · https://help.aliyun.com/en/model-studio/qwen-audio-realtime-user-guides · https://arxiv.org/abs/2609.25176","page_url":"https://postcutoff.com/m/qwen-audio-3-1-realtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":322,"you_url":null,"briefings":null},{"id":"claude-opus-5-5","name":"Claude Opus 5.5","org":"Anthropic","family":"Claude 5","released":"2026-09-22","status":"current","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-06","pricing":{"input":4,"output":20,"cache_read":0.2,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$4 in, $20 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-5-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-5-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-opus-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-5-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-5.5","url":"https://openrouter.ai/anthropic/claude-opus-5.5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Top agentic coding at lower cost","detail":"Anthropic reports 66.4% on Terminal-Bench 4.0, ahead of GPT-6 Astra at roughly 40% of the cost; an early tester finished a 680k-line code migration in under a day.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-opus-5-5"},{"name":"Knowledge-work lead (GDPval-AA)","detail":"Launch claim of 1846 Elo on GDPval-AA v2.1, above both Claude Fable 5.1 (1735) and Claude Opus 5 (1708).","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-opus-5-5"},{"name":"Cheaper, faster Opus","detail":"About 40% cheaper than Opus 5 on typical workloads ($4/$20 per MTok, cache reads $0.20) and about 30% faster output at default settings.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-opus-5-5"},{"name":"Opus with Fable-level safeguards","detail":"Anthropic says it is the first Opus model whose safeguards match Claude Fable 5.1 on cyber, bio and distillation (refusal categories include bio and reasoning_extraction).","first":true,"discovered":"launch","source":"https://www.anthropic.com/claude-opus-5-5"},{"name":"Thinking that cannot be disabled","detail":"Adaptive thinking is always on and effort is the only control (default medium). Text between tool calls comes back as progress-update thinking blocks.","first":false,"discovered":"launch","source":"https://platform.claude.com/docs/en/models/opus-5-5/overview"}],"entry":null,"notes":"Anthropic's recommended default model. Thinking always on (cannot be disabled); effort default is medium (set explicitly); forced tool_choice any/tool returns 400; computer use only via computer_toolset_20260801 on Claude API/Google Cloud. Fast mode (Claude API only) $8/$40. Batch $2/$10; up to 300K output on Batch with output-300k-2026-03-24 beta.","verified":"2026-09-29","body_md":"Default choice for most workloads: long-running agentic coding and knowledge work, cheaper than Claude Opus 5 ($5/$25). Text+image in, text out.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-5-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-5-5/overview\n- Announcement: https://www.anthropic.com/claude-opus-5-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-5-5/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":302,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-5-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"gemini-3-8-flash-tts","name":"Gemini 3.8 Flash TTS","org":"Google DeepMind","family":"Gemini 3","released":"2026-09-22","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":8192,"max_output":16384,"knowledge_cutoff":null,"pricing":{"input":0.5,"output":9,"unit":"per 1M tokens (text in / audio out; introductory through 2026-12-31, $1.00 / $18.00 from 2027-01-01)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.50 in, $9 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.8-flash-tts","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash-tts:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts"},{"provider":"Gemini API","model_id":"gemini-3.8-flash-lite-tts","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-lite-tts"}],"capabilities":[{"name":"Voice design from prompts","detail":"Create entirely new voices from natural-language descriptions; #1 on Hume AI Voice Design Benchmark.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/"},{"name":"#1 on Hume Real-World VoiceEQ leaderboard","detail":"Hume's blind human-rated benchmark (2026-09-24): Gemini 3.8 Flash TTS 0.920 and Flash-Lite TTS 0.914 expressivity-reliability score, ahead of Gemini 2.5 Pro TTS (0.880) and Cartesia Sonic 3.6 (0.840); long-form stability up from 1.22 to ~2.9-3.0/5, but weaker speaker similarity (3.68/5).","first":false,"discovered":"later","source":"https://www.hume.ai/blog/newly-released-google-s-gemini-3-8-flash-tts-tops-hume-s-real-world-voiceeq-leaderboard"},{"name":"Voice replication","detail":"Recreates a consistent voice from a ~30-second sample with consent verification; 2,000+ library voices.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/"},{"name":"Directed long-form multi-speaker audio","detail":"Line-by-line direction of pacing/emotion, dual-speaker staging, stable over hours; 130+ languages with auto-detection.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts"}],"entry":null,"notes":"Sibling gemini-3.8-flash-lite-tts (101 languages) costs $0.50 in / $6.00 audio out (intro). Outputs SynthID-watermarked. Older: gemini-3.1-flash-tts-preview, gemini-2.5-flash-preview-tts, gemini-2.5-pro-preview-tts.","verified":"2026-09-29","body_md":"Studio-grade, steerable text-to-speech (audiobooks, games, dubbing, voice agents).\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash-tts:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Say cheerfully: Have a wonderful day!\"}]}],\n       \"generationConfig\":{\"responseModalities\":[\"AUDIO\"]}}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/).","page_url":"https://postcutoff.com/m/gemini-3-8-flash-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":343,"you_url":null,"briefings":null},{"id":"gpt-6-luna","name":"GPT-6 Luna","org":"OpenAI","family":"GPT-6","released":"2026-09-22","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-05","pricing":{"input":0.1,"cached_input":0.01,"output":0.5,"cache_write":0.125,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.10 in, $0.50 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-6-luna","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-6-luna"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-6-luna","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-6-luna","url":"https://openrouter.ai/openai/gpt-6-luna"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"1M context at $0.10/M","detail":"Cheapest OpenAI reasoning model with the full 1.05M context window and 128K output.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6-luna"},{"name":"Agentic tools on the budget tier","detail":"Supports computer use, hosted shell, MCP and tool search like the larger models.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6-luna"},{"name":"Free-tier ChatGPT model","detail":"Rolled out to ChatGPT free users and the desktop app at launch.","first":false,"discovered":"launch","source":"https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/"},{"name":"Decisions API model","detail":"Only model behind OpenAI's Decisions API (POST /v1/decisions, public beta since 2026-10-06): typed predicate/choice/score answers from text and images, about 10x faster than the Responses API, billed at $0.10 per 1M input tokens with no output charge.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/guides/decisions"}],"entry":null,"notes":"Most efficient GPT-6 model for focused, high-volume tasks; successor to GPT-5.6 Luna (the mini/nano tier). OpenRouter also lists openai/gpt-6-luna-pro.","verified":"2026-10-07","body_md":"High-volume classification, extraction, routing, sub-agents.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-6-luna\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-6-luna\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-6-luna/","events_after":694,"major_after":173,"historic_after":37,"missing_at_launch":331,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-6-luna/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-05-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-05-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-05-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-05.md"}},{"id":"gpt-6-sol","name":"GPT-6 Sol","org":"OpenAI","family":"GPT-6","released":"2026-09-22","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-04","pricing":{"input":2,"cached_input":0.2,"output":10,"cache_write":2.5,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2 in, $10 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-6-sol","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-6-sol"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-6-sol","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-6-sol","url":"https://openrouter.ai/openai/gpt-6-sol"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Astra-level reliability at lower cost","detail":"OpenAI claims about half as many mistakes as GPT-5.6 Sol at half its API price.","first":false,"discovered":"launch","source":"https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/"},{"name":"Full hosted tool suite","detail":"Web/file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6-sol"},{"name":"Image-input bug fix","detail":"Sep 25 2026 fix for an image-encoding bug that degraded image understanding at launch.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/changelog"}],"entry":null,"notes":"Superseded on Sept 29 2026 by GPT-6.1 Sol (same $2/$10 price, cached input $0.10; see gpt-6-1-sol) but still listed on the pricing page. Mid-tier GPT-6 model for complex coding and agentic workflows; successor to GPT-5.6 Sol. Reasoning effort none..max. OpenRouter also lists openai/gpt-6-sol-pro (reasoning.mode pro). ARC Prize verified (Sept 28 2026) - ARC-AGI-3 4.6% standard harness / 23.0% provider harness, ARC-AGI-2 89.6% at $0.44/task (x.com/arcprize/status/2104671274299429283).","verified":"2026-09-29","body_md":"Default choice for coding agents and complex workflows at moderate cost.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-6-sol\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-6-sol\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-6-sol/","events_after":725,"major_after":180,"historic_after":38,"missing_at_launch":362,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-6-sol/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-04-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-04-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-04-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-04.md"}},{"id":"qwen-audio-3-1-tts-next","name":"Qwen-Audio-3.1-TTS-Next","org":"Alibaba (Qwen)","family":"Qwen-Audio 3.1","released":"2026-09-22","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.848,"output":1.696,"unit":"USD per 1M tokens (China/Beijing region price shown in docs; international price not listed)","source":"https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next"},"price_line":"$0.848 in, $1.696 out per 1M tokens","access":[{"provider":"Alibaba Cloud Model Studio","model_id":"qwen-audio-3.1-tts-next","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next"}],"capabilities":[{"name":"One-pass speech + sound effects + ambience","detail":"'AudioGen' model (LM + diffusion) that generates complete audio - speech, multi-speaker dialogue, podcasts, sound effects and ambient soundscapes - in a single pass from text, timestamps and up to 3 reference clips.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next"}],"entry":"2026-09-23-qwen-audio-3-1","notes":"Chinese and English only; max 3,000 input characters; output up to 240 s for podcasts, 120 s otherwise. Comparable to ByteDance Seed Audio 1.0 (Jul 2026) and StepAudio 3 Gen. Sibling TTS model Qwen-Audio-3.1-TTS (plain TTS, ~70% cheaper than 3.0) exists but its exact API id was not verified: as of 2026-09-29 the international Model Studio docs (models page, qwen-tts page) list only qwen-audio-3.0-tts-flash / -plus, and neither qwen-audio-3.1-tts-flash/-plus nor an ASR-Next id resolves on QwenCloud (404). Verified 3.1 ASR ids: qwen-audio-3.1-asr-flash(-streaming/-filetrans).","verified":"2026-09-29","body_md":"Scene-level audio creation (audiobooks, podcasts, games, ads) rather than plain TTS.\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen-audio-3-1-tts-next","page_url":"https://postcutoff.com/m/qwen-audio-3-1-tts-next/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":343,"you_url":null,"briefings":null},{"id":"grok-4-7","name":"Grok 4.7","org":"xAI","family":"Grok 4","released":"2026-09-21","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":500000,"max_output":null,"knowledge_cutoff":"2026-05","pricing":{"input":2,"cached_input":0.5,"output":6,"input_over_200k":4,"output_over_200k":12,"unit":"per 1M tokens (USD); higher tier applies to whole request when prompt >= 200k tokens","source":"https://docs.x.ai/docs/models"},"price_line":"$2 in, $6 out per 1M tokens","access":[{"provider":"xAI API","model_id":"grok-4.7","endpoint":"https://api.x.ai/v1/chat/completions","docs":"https://docs.x.ai/docs/models/grok-4.7"},{"provider":"AWS Bedrock","model_id":"xai.grok-4.7","inference_profiles":["global.xai.grok-4.7","us.xai.grok-4.7"],"docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-7.html"},{"provider":"OpenRouter","model_id":"x-ai/grok-4.7","url":"https://openrouter.ai/x-ai/grok-4.7"},{"provider":"Web app","url":"https://grok.com"}],"capabilities":[{"name":"Four-level reasoning effort incl. xhigh","detail":"Configurable reasoning effort low / medium / high / xhigh (default high) on one model id.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models/grok-4.7"},{"name":"500K context at unchanged price","detail":"500K-token context with text+image input; launched at the same $2/$6 price as Grok 4.6 while claiming notable gains.","first":false,"discovered":"launch","source":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-7.html"},{"name":"Mixed independent benchmark results","detail":"Early third-party evals showed a more mixed picture than xAI's claims, still behind top Claude/GPT-6 models on several tasks.","first":false,"discovered":"later","source":"https://tech.yahoo.com/ai/gemini/articles/xai-launches-grok-4-7-171603280.html"}],"entry":null,"notes":"Alias grok-4.7-latest. xAI flagship as of Sept 2026; no Batch API; logprobs unsupported. Bedrock launched 2026-09-28 (Global CRIS $2/$6, Geo $2.20/$6.60). Max output not published.","verified":"2026-09-29","body_md":"xAI's flagship for coding, agentic tasks and knowledge work (successor to Grok 4.6 / 4.5).\n\n```bash\ncurl https://api.x.ai/v1/chat/completions -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-4.7\",\"reasoning_effort\":\"high\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.x.ai/docs/models , https://docs.x.ai/docs/models/grok-4.7 , https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-7.html","page_url":"https://postcutoff.com/m/grok-4-7/","events_after":694,"major_after":173,"historic_after":37,"missing_at_launch":314,"events_since_release":null,"you_url":"https://postcutoff.com/you/grok-4-7/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-05-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-05-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-05-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-05.md"}},{"id":"mimo-v2-6-flash","name":"MiMo-V2.6-Flash","org":"Xiaomi","family":"MiMo V2.6","released":"2026-09-21","status":"current","type":"reasoning-llm","modality_in":["text","image","video","audio"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":1000000,"max_output":128000,"knowledge_cutoff":null,"pricing":{"input":0.14,"cached_input":0.0028,"output":0.28,"unit":"per 1M tokens (USD)","source":"https://venturebeat.com/technology/better-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash"},"price_line":"$0.14 in, $0.28 out per 1M tokens","access":[{"provider":"Xiaomi MiMo API","model_id":"mimo-v2.6-flash","docs":"https://mimo.mi.com/"},{"provider":"Hugging Face","url":"https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL"}],"capabilities":[{"name":"Cheap open omnimodal MoE","detail":"~311B total / 15B active, 1M context, MIT license, at $0.14 / $0.28 per 1M tokens; RL post-training reportedly cost ~$850K.","first":false,"discovered":"launch","source":"https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL"}],"entry":"2026-09-21-xiaomi-mimo-v2-6","notes":"A 9B distill (MiMo-V2.6-Distill-Qwen-9B) was released alongside.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/mimo-v2-6-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":363,"you_url":"https://postcutoff.com/you/mimo-v2-6-flash/","briefings":null},{"id":"mimo-v2-6-pro","name":"MiMo-V2.6-Pro","org":"Xiaomi","family":"MiMo V2.6","released":"2026-09-21","status":"current","type":"reasoning-llm","modality_in":["text","image","video","audio"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":1000000,"max_output":128000,"knowledge_cutoff":null,"pricing":{"input":0.435,"cached_input":0.0036,"output":0.87,"unit":"per 1M tokens (USD; list price ¥3 / ¥6)","source":"https://mimo.mi.com/models/en-US/mimo-v2.6-pro"},"price_line":"$0.435 in, $0.87 out per 1M tokens","access":[{"provider":"Xiaomi MiMo API","model_id":"mimo-v2.6-pro","docs":"https://mimo.mi.com/models/en-US/mimo-v2.6-pro"},{"provider":"Hugging Face","url":"https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL"},{"provider":"OpenRouter (Pro-UltraSpeed tier, ~20x output speed, $4.35 / $8.70)","model_id":"see https://openrouter.ai (Xiaomi MiMo V2.6 Pro UltraSpeed)"}],"capabilities":[{"name":"Top open-weights model on the Artificial Analysis Intelligence Index","detail":"Scored 46 at launch (Sept 2026), tying Grok 4.7 and ahead of DeepSeek V4.1 Flash (39).","first":false,"discovered":"launch","source":"https://venturebeat.com/technology/better-than-deepseek-xiaomis-mimo-v2-6-pro-debuts-as-the-top-open-weights-model-in-the-world-alongside-cheaper-v2-6-flash"},{"name":"Omnimodal 1T-parameter MIT-licensed MoE","detail":"1.02T total / 42B active, 70 layers (60 SWA + 10 global), 1M context; text, image, video and audio input under MIT.","first":false,"discovered":"launch","source":"https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL"}],"entry":"2026-09-21-xiaomi-mimo-v2-6","notes":"Xiaomi-reported: DeepSWE v1.1 71.9, Terminal-Bench 2.1 89.9, AutomationBench 53.1, CyberGym 94.0. MOPD (distilled) variant added ~Sept 27.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/mimo-v2-6-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":363,"you_url":"https://postcutoff.com/you/mimo-v2-6-pro/","briefings":null},{"id":"qwen-image-2-1","name":"Qwen-Image-2.1","org":"Alibaba (Qwen)","family":"Qwen-Image","released":"2026-09-20","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":true,"model_license":"qwen-research","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"Qwen/Qwen-Image-2.1","url":"https://huggingface.co/Qwen/Qwen-Image-2.1"},{"provider":"ModelScope","url":"https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1"},{"provider":"GitHub","url":"https://github.com/QwenLM/Qwen-Image-2.1"},{"provider":"Hugging Face Space (demo)","url":"https://huggingface.co/spaces/Qwen/Qwen-Image-2.1"}],"capabilities":[{"name":"Native RGBA generation and editing","detail":"Generates transparent images, edits transparent layers and extracts subjects from photos in one model.","first":false,"discovered":"launch","source":"https://github.com/QwenLM/Qwen-Image-2.1"},{"name":"Multi-reference editing","detail":"Up to 10 reference images; local edits via circles, painted annotations or masks, with identity preservation.","first":false,"discovered":"launch","source":"https://github.com/QwenLM/Qwen-Image-2.1"}],"entry":"2026-09-20-qwen-image-2-1","notes":"7B visual generator (32 single-stream DiT layers). Diffusers pipeline QwenImage21Pipeline (install diffusers from git); day-0 ComfyUI, vLLM-Omni, SGLang support. Qwen Research License, check terms before commercial use.","verified":"2026-09-30","body_md":"Quick start (from the README):\n\n```python\nimport torch\nfrom diffusers import QwenImage21Pipeline\npipe = QwenImage21Pipeline.from_pretrained(\"Qwen/Qwen-Image-2.1\", torch_dtype=torch.bfloat16).to(\"cuda\")\nimage = pipe(prompt=\"A neon shop sign that reads \\\"QWEN IMAGE 2.1\\\"\", num_inference_steps=40).images[0]\n```","page_url":"https://postcutoff.com/m/qwen-image-2-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":380,"you_url":null,"briefings":null},{"id":"step-5-preview","name":"Step 5 Preview","org":"StepFun","family":"Step 5","released":"2026-09-20","status":"preview","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":64000,"knowledge_cutoff":null,"pricing":{"input":1,"cached_input":0.05,"output":2.7,"unit":"per 1M tokens (USD); reasoning tokens billed as output","source":"https://pandaily.com/stepfun-step-5-preview-600b-moe-1m-context-open-weights-oct-15"},"price_line":"$1 in, $2.70 out per 1M tokens","access":[{"provider":"StepFun API","model_id":"step-5-preview","docs":"https://platform.stepfun.ai/docs/en/guides/models/step-5-preview"}],"capabilities":[{"name":"600B sparse MoE agent model with 1M context and video input","detail":"600B total / 27B active, 92 layers; up to 60 images and short videos per request; reasoning effort low/medium/high.","first":false,"discovered":"launch","source":"https://platform.stepfun.ai/docs/en/guides/models/step-5-preview"}],"entry":"2026-09-20-stepfun-step-5-preview","notes":"Open weights promised for 2026-10-15 (placeholder HF repo stepfun-ai/Step-5-Preview-BF16); update open_weights/license then. Pricing from press, not verified on StepFun's pricing page.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/step-5-preview/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":380,"you_url":"https://postcutoff.com/you/step-5-preview/","briefings":null},{"id":"qwen3-8-livetranslate","name":"Qwen3.8-LiveTranslate (Flash Realtime)","org":"Alibaba (Qwen)","family":"Qwen3.8","released":"2026-09-19","status":"current","type":"audio/speech","modality_in":["audio","image","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":53248,"max_output":4096,"knowledge_cutoff":null,"pricing":{"audio_input":7.5,"image_input":0.55,"text_output":20,"audio_output":30,"unit":"USD per 1M tokens (QwenCloud list price; press estimates about $1.54 per hour of speech in and out)","source":"https://www.qwencloud.com/models/qwen3.8-livetranslate-flash-realtime"},"price_line":"$7.50 audio in, $20 text out per 1M tokens","access":[{"provider":"QwenCloud (Realtime WebSocket)","model_id":"qwen3.8-livetranslate-flash-realtime","endpoint":"wss://maas.qwencloudapi.com/api-ws/v1/realtime?model=qwen3.8-livetranslate-flash-realtime","docs":"https://www.qwencloud.com/models/qwen3.8-livetranslate-flash-realtime"},{"provider":"Alibaba Cloud Model Studio","model_id":"qwen3.8-livetranslate-flash-realtime","docs":"https://www.alibabacloud.com/help/en/model-studio/models"}],"capabilities":[{"name":"Simultaneous interpretation with lower lag","detail":"Streams translated speech and text while the speaker is still talking; average lagging (LAAL) cut from 2.8 s to 2.3 s across 60 languages with a new 'Interleave' architecture.","first":false,"discovered":"launch","source":"https://x.com/Alibaba_Qwen/status/2101206705111757253"},{"name":"Multi-speaker diarization with per-speaker voice cloning","detail":"Tells speakers apart in multi-party speech and keeps each speaker's own voice in the translated audio; synchronized bilingual on-screen display.","first":false,"discovered":"launch","source":"https://x.com/Alibaba_Qwen/status/2101206705111757253"},{"name":"Long-context disambiguation","detail":"Uses conversation history to keep names and terminology consistent across a session.","first":false,"discovered":"launch","source":"https://x.com/Alibaba_Qwen/status/2101206705111757253"}],"entry":"2026-09-23-qwen-audio-3-1","notes":"Understands 60 languages and speaks 29 (the rest get text-only translation). Thinker-talker hybrid MoE on the Qwen-Omni stack (press). API-only, no open weights and no announced timeline for them. MindStudio's hands-on found short sentences fine but weak end-of-turn detection, so developers need their own turn-taking logic. Announced on X 2026-09-19 (294k views by 2026-09-29), shortly before Apsara 2026.","verified":"2026-09-29","body_md":"Alibaba's hosted real-time interpretation model, a competitor to gpt-realtime-translate and Gemini 3.5 Live Translate.\n\nSources: https://x.com/Alibaba_Qwen/status/2101206705111757253 · https://qwen.ai/blog?id=qwen3.8-livetranslate · https://www.qwencloud.com/models/qwen3.8-livetranslate-flash-realtime · https://www.marktechpost.com/2026/09/19/alibaba-qwen-team-releases-qwen3-8-livetranslate/ · https://www.mindstudio.ai/blog/qwen3-8-livetranslate-hands-on","page_url":"https://postcutoff.com/m/qwen3-8-livetranslate/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":383,"you_url":null,"briefings":null},{"id":"grok-voice-transcribe-2-0","name":"Grok Voice Transcribe 2.0","org":"xAI","family":"Grok Voice","released":"2026-09-18","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour_batch":0.1,"per_hour_streaming":0.2,"unit":"USD per hour of audio (REST batch / WebSocket streaming); diarization, timestamps and key-term biasing included","source":"https://x.ai/news/grok-voice-transcribe-2"},"price_line":"$0.10 per hour, batch","access":[{"provider":"xAI API (REST)","model_id":"grok-voice-transcribe-2.0","endpoint":"https://api.x.ai/v1/stt","docs":"https://docs.x.ai/developers/model-capabilities/audio/speech-to-text"},{"provider":"xAI API (WebSocket streaming)","model_id":"grok-voice-transcribe-2.0","endpoint":"wss://api.x.ai/v1/stt","docs":"https://docs.x.ai/developers/model-capabilities/audio/speech-to-text"}],"capabilities":[{"name":"Top streaming STT accuracy (claimed)","detail":"xAI says it ranks #1 for accuracy among 32 streaming models on the Artificial Analysis leaderboard; multilingual short-phrase WER 20.6% -> 6.8% vs v1.0 ('2x as accurate').","first":false,"discovered":"launch","source":"https://x.ai/news/grok-voice-transcribe-2"},{"name":"Very low price with diarization included","detail":"$0.10/hr batch and $0.20/hr streaming, with speaker diarization, word timestamps, up to 8-channel multichannel and 100 key terms per request at no extra cost.","first":false,"discovered":"launch","source":"https://x.ai/news/grok-voice-transcribe-2"}],"entry":null,"notes":"Launched 2026-09-18 as drop-in upgrade of the Grok STT API (first released 2026-04-17 with grok-voice-transcribe-1.0, which can be pinned but will be deprecated). Up to 500 MB files; WAV/MP3/OGG/Opus/FLAC/AAC/MP4/M4A/MKV plus raw PCM/mu-law/A-law at 8-48 kHz; Smart Turn end-of-turn detection, VAD, inverse text normalization, filler removal, mid-recording language switching. Docs list ~25 languages for formatting.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.x.ai/v1/stt -H \"Authorization: Bearer $XAI_API_KEY\" \\\n  -F model=grok-voice-transcribe-2.0 -F file=@call.wav -F diarize=true\n```\n\nSources: https://x.ai/news/grok-voice-transcribe-2 · https://docs.x.ai/developers/model-capabilities/audio/speech-to-text · https://x.ai/news/grok-stt-and-tts-apis","page_url":"https://postcutoff.com/m/grok-voice-transcribe-2-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":386,"you_url":null,"briefings":null},{"id":"figure-helix-2-5","name":"Helix 2.5","org":"Figure AI","family":"Helix","released":"2026-09-17","status":"current","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"None (runs only on Figure 03 robots; no public API, weights or waitlist)","url":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"}],"capabilities":[{"name":"Zero-shot whole-body generalization to unseen homes","detail":"56% success (237/420 trials) tidying, towel folding and bed making in 30 never-seen Bay Area homes with no data from those homes; matched Helix 02's success with half the adaptation data.","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"},{"name":"Pretrained from scratch on human video (Index)","detail":"Pretrained from random initialization on Figure's Index human-video dataset (not a VLM); without it the same model scored 9%.","first":false,"discovered":"launch","source":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"},{"name":"Human-to-robot transfer scaling law","detail":"Predictable scaling of robot performance with human-video data (forecast error 0.54% over an 8x data range).","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization"}],"entry":"2026-09-17-figure-helix-2-5","notes":"'first' flags are Figure's 'to our knowledge' claims (first zero-shot whole-body generalization at this scope; first human-to-robot transfer scaling law measured on a humanoid). Company-reported results. Architecture/parameter counts not disclosed.","verified":null,"body_md":"Sources: [Figure: Helix 2.5](https://www.figure.ai/news/helix-2-5-zero-shot-30-home-generalization).","page_url":"https://postcutoff.com/m/figure-helix-2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":397,"you_url":null,"briefings":null},{"id":"speechmatics-linden-1","name":"Speechmatics Linden 1 (Agent STT)","org":"Speechmatics","family":"Speechmatics STT","released":"2026-09-17","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.3,"unit":"USD per audio hour launch pricing; $0.16/hour with volume discount","source":"https://www.globenewswire.com/news-release/2026/09/17/3364138/0/en/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents.html"},"price_line":"$0.30 per hour","access":[{"provider":"Speechmatics Agent STT API","model_id":"linden-1","endpoint":"/v2/agent (regions eu1 / us1 .asr.api.speechmatics.com)","docs":"https://docs.speechmatics.com/speech-to-text/models"},{"provider":"Pipecat","url":"https://www.speechmatics.com/voice-agents"},{"provider":"LiveKit","url":"https://docs.livekit.io/agents/models/stt/speechmatics/"}],"capabilities":[{"name":"STT output shaped for LLM voice agents","detail":"Returns speaker-attributed segments with turn messages instead of a running word stream; finalizes segments in under 350 ms; 55+ languages; custom vocabulary up to 1,000 terms; live diarization and speaker ID.","first":false,"discovered":"launch","source":"https://docs.speechmatics.com/speech-to-text/models"},{"name":"Low semantic error on Pipecat benchmark","detail":"1.05% pooled semantic error rate and 369 ms median finalization on the Pipecat STT benchmark (23 streaming models), on the speed/accuracy Pareto frontier, per Speechmatics.","first":false,"discovered":"launch","source":"https://www.globenewswire.com/news-release/2026/09/17/3364138/0/en/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents.html"}],"entry":"2026-09-17-speechmatics-agent-stt-linden","notes":"Targets high-consequence errors in calls (a changed digit, a missed 'not', a one-word confirmation). Benchmark figures are vendor-reported from Pipecat's public benchmark. Sibling batch model: speechmatics-melia-1.","verified":"2026-09-29","body_md":"Sources: https://www.globenewswire.com/news-release/2026/09/17/3364138/0/en/speechmatics-launches-agent-stt-for-the-speech-errors-that-derail-voice-agents.html , https://docs.speechmatics.com/speech-to-text/models","page_url":"https://postcutoff.com/m/speechmatics-linden-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":397,"you_url":null,"briefings":null},{"id":"gemini-3-8-live","name":"Gemini 3.8 Live","org":"Google DeepMind","family":"Gemini 3","released":"2026-09-15","status":"current","type":"audio/speech","modality_in":["text","image","audio","video"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":0.75,"output":4.5,"audio_input":3,"audio_output":12,"unit":"per 1M tokens (text in $0.75, text out $4.50; audio in $3.00 = ~$0.005/min, audio out $12.00 = ~$0.018/min)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.75 in, $4.50 out per 1M tokens","access":[{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-3.8-live","endpoint":"wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live"},{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-3.8-live-extended-thinking","docs":"https://ai.google.dev/gemini-api/docs/live-api"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Real-time multilingual voice agents","detail":"Low-latency speech-to-speech with near-real-time visual grounding; 97 languages with mid-conversation switching.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/"},{"name":"Asynchronous tool use while talking","detail":"Keeps the conversation going while tools run in the background, narrating progress ('Let me check that...').","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/"},{"name":"Extended Thinking variant tops S2S quality","detail":"gemini-3.8-live-extended-thinking ranked #1 on Artificial Analysis Speech-to-Speech Quality Index (82.6) and 97.7% Big Bench Audio.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/"}],"entry":null,"notes":"Default Live API model; thinking_level not supported on gemini-3.8-live (use gemini-3.8-live-extended-thinking for deeper reasoning; pricing page lists it at the same rates as gemini-3.8-live, checked 2026-09-29). Previous: gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-preview-12-2025. WebSocket endpoint is the standard Live API URL, not re-read today.","verified":"2026-09-29","body_md":"Native-audio model for real-time voice and video agents via the Live API (WebSocket, bidirectional streaming).\n\n```python\nfrom google import genai\nclient = genai.Client()\nasync with client.aio.live.connect(model=\"gemini-3.8-live\",\n        config={\"response_modalities\": [\"AUDIO\"]}) as session:\n    await session.send_client_content(turns={\"parts\": [{\"text\": \"Hi!\"}]})\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/).","page_url":"https://postcutoff.com/m/gemini-3-8-live/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":423,"you_url":null,"briefings":null},{"id":"stepaudio-3-asr-tts","name":"StepAudio 3 ASR Max / StepAudio 3 TTS","org":"StepFun","family":"StepAudio 3","released":"2026-09-15","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour_asr":0.4,"per_10k_characters_tts":0.36,"unit":"USD: stepaudio-3-asr-max $0.40/hour of audio; stepaudio-3-tts $0.36 per 10,000 characters; voice cloning $1.50/voice","source":"https://platform.stepfun.ai/docs/en/pricing/details"},"price_line":"$0.40 per hour, speech recognition","access":[{"provider":"StepFun API","model_id":"stepaudio-3-asr-max","docs":"https://platform.stepfun.ai/docs/en/guides/models/audio"},{"provider":"StepFun API","model_id":"stepaudio-3-tts","docs":"https://platform.stepfun.ai/docs/en/guides/models/audio"},{"provider":"StepFun API (preview)","model_id":"stepaudio-3-gen-preview","docs":"https://platform.stepfun.ai/docs/en/guides/models/audio"}],"capabilities":[{"name":"#1 non-streaming ASR on AA-WER","detail":"Artificial Analysis ranked StepAudio 3 ASR #1 on its AA-WER Index for non-streaming speech-to-text with 1.7% WER (StepAudio 2.5 ASR: 4.7%).","first":false,"discovered":"launch","source":"https://x.com/ArtificialAnlys/status/2102485740248842710"},{"name":"Context-aware streaming TTS","detail":"Natural, context-aware speech with low-latency streaming, natural-language control and voice cloning; 1,000-char input limit; wav/mp3/flac/opus/pcm.","first":false,"discovered":"launch","source":"https://platform.stepfun.ai/docs/en/guides/models/audio"}],"entry":"2026-09-15-stepfun-stepaudio-3","notes":"Family file for the non-realtime StepAudio 3 models. Languages: zh, en, ja, ko, fr, es (non-zh/en in preview). stepaudio-3-gen-preview (speech+SFX+ambience+BGM) and stepaudio-3-music-preview are free during preview. Previous gen: stepaudio-2.5-asr ($0.022/h), stepaudio-2.5-asr-stream ($0.18/h), stepaudio-2.5-tts ($0.85/10k chars). Exact per-model release date assumed = family launch 2026-09-15.","verified":"2026-09-29","body_md":"Sources: https://platform.stepfun.ai/docs/en/guides/models/audio · https://platform.stepfun.ai/docs/en/pricing/details","page_url":"https://postcutoff.com/m/stepaudio-3-asr-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":423,"you_url":null,"briefings":null},{"id":"stepaudio-3-realtime","name":"StepAudio 3 Realtime","org":"StepFun","family":"StepAudio 3","released":"2026-09-15","status":"preview","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0,"output":0,"unit":"free during limited-time preview (successor stepaudio-2.5-realtime: $1.50 in / $0.30 cached / $10.00 out per 1M tokens)","source":"https://platform.stepfun.ai/docs/en/pricing/details"},"price_line":"Free","access":[{"provider":"StepFun API (Realtime WebSocket)","model_id":"stepaudio-3-realtime-preview","docs":"https://platform.stepfun.ai/docs/en/guides/models/stepaudio-3-realtime"},{"provider":"StepFun API (Chat Completions)","model_id":"stepaudio-3-chat-preview","docs":"https://platform.stepfun.ai/docs/en/guides/models/stepaudio-3-realtime"}],"capabilities":[{"name":"Think-while-speaking full duplex","detail":"Runs private chain-of-thought in parallel with spoken output; distinguishes real interruptions from backchannels; asynchronous tool execution (web search, knowledge retrieval).","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2609.14005"},{"name":"#1 on Artificial Analysis conversational dynamics","detail":"98.9 on Artificial Analysis Full-Duplex Bench (Conversational Dynamics) and 99.7% Speech Reasoning at launch, ahead of Qwen Audio 3.0 Realtime Plus and GPT-Live-1 per StepFun.","first":false,"discovered":"launch","source":"https://x.com/StepFun_ai/status/2099916376274313630"}],"entry":"2026-09-15-stepfun-stepaudio-3","notes":"Chinese and English. Preview ids will be retired for a paid GA version when the trial ends. Predecessor stepaudio-2.5-realtime (2026-05-26; persona role-play, project page https://stepaudiollm.github.io/step-audio-2.5-realtime/ with self-reported 86.36 general dialogue / 79.80 spoken QA / 82.18 paralinguistics, claimed to beat GPT-Realtime-1.5 on StepFun's evals). Technical report arXiv 2609.14005 (56.0% task success on tau-Voice).","verified":"2026-09-29","body_md":"StepFun's full-duplex voice agent model; part of the five-model StepAudio 3 family (Realtime, ASR Max, TTS, Gen, Music).\n\nSources: https://platform.stepfun.ai/docs/en/guides/models/audio · https://platform.stepfun.ai/docs/en/pricing/details · https://arxiv.org/abs/2609.14005","page_url":"https://postcutoff.com/m/stepaudio-3-realtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":423,"you_url":null,"briefings":null},{"id":"elevenlabs-music-v2-5","name":"Eleven Music v2.5","org":"ElevenLabs","family":"Eleven Music","released":"2026-09-11","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.15,"unit":"per minute of generated music (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.15 per minute","access":[{"provider":"ElevenLabs API","model_id":"music_v2_5","endpoint":"https://api.elevenlabs.io/v1/music","docs":"https://elevenlabs.io/docs/overview/capabilities/music"},{"provider":"Web app (ElevenMusic)","url":"https://elevenmusic.io"}],"capabilities":[{"name":"Commercially cleared music generation","detail":"Richer melodies and live-sounding instruments, built for commercial use; lossless downloads on every plan incl. Free.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/music-v2-5-model"},{"name":"Composition plans and audio reference","detail":"Music v2 line supports structured composition plans and reference-audio generation (v2.5 default for prompted and reference generation).","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"},{"name":"Composition-plan chunks via API","detail":"API support rolled out 2026-09-14 with 6,132-character composition chunks; waveform visual data via with_waveform_visual (2026-08-03).","first":false,"discovered":"later","source":"https://elevenlabs.io/docs/changelog"}],"entry":"2026-09-11-elevenlabs-music-v2-5","notes":"Announced 2026-09-11 (blog + YouTube). music_v2 and music_v1 remain available (v1 'outclassed by v2/v2.5'). Preferred over v2 in a blind test of 47,885 sample pairs; biggest gains in R&B/soul, hip hop/trap, rock/metal, orchestral/cinematic. Downloads: Free 5 lossless/day, Pro 400/month; tracks based on other artists' songs cannot be downloaded (protections built with labels/publishers).","verified":"2026-09-29","body_md":"Text-to-music (vocals or instrumental) via API or ElevenMusic.\n\n```bash\ncurl -X POST https://api.elevenlabs.io/v1/music -H \"xi-api-key: $ELEVENLABS_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"prompt\":\"upbeat indie pop, female vocals, summer road trip\",\"music_length_ms\":30000,\"model_id\":\"music_v2_5\"}' -o song.mp3\n```\n\nLaunch video: https://www.youtube.com/watch?v=zXlVQ8rMJM0\n\nSources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/blog/music-v2-5-model , https://elevenlabs.io/docs/changelog","page_url":"https://postcutoff.com/m/elevenlabs-music-v2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":440,"you_url":null,"briefings":null},{"id":"deepseek-v4-1-flash","name":"DeepSeek-V4.1-Flash","org":"DeepSeek","family":"DeepSeek V4","released":"2026-09-10","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":1000000,"max_output":384000,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":1.2,"cache_read":0.006,"unit":"per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.15, output 0.6, cache hit 0.003). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri","source":"https://api-docs.deepseek.com/quick_start/pricing"},"price_line":"$0.30 in, $1.20 out per 1M tokens","access":[{"provider":"DeepSeek API","model_id":"deepseek-flash","endpoint":"https://api.deepseek.com","docs":"https://api-docs.deepseek.com/quick_start/pricing"},{"provider":"DeepSeek API (Anthropic format)","model_id":"deepseek-flash","endpoint":"https://api.deepseek.com/anthropic","docs":"https://api-docs.deepseek.com/quick_start/pricing"},{"provider":"Alibaba Cloud Model Studio","model_id":"deepseek-v4.1-flash","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"OpenRouter","model_id":"deepseek/deepseek-v4.1-flash","url":"https://openrouter.ai/deepseek/deepseek-v4.1-flash"},{"provider":"Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"},{"provider":"Web app","url":"https://chat.deepseek.com"}],"capabilities":[{"name":"Native vision in the Flash tier","detail":"First DeepSeek Flash model with native multimodal (image) understanding built in; replaced the separate V4-Flash-Vision-Exp.","first":false,"discovered":"launch","source":"https://api-docs.deepseek.com/updates"},{"name":"Causal Encoder-Decoder (CED) architecture","detail":"552B-backbone MoE that activates only ~8B params per token in prefill and ~16B in decode, aimed at input-heavy agentic workloads.","first":false,"discovered":"launch","source":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"},{"name":"Tiny KV cache (CSA2 + FP4 KV)","detail":"Compressed Sparse Attention 2 and FP4 main KV cache cut the global KV cache to ~890 bytes/token, about 1/4 of V4-Flash.","first":false,"discovered":"launch","source":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"},{"name":"Hybrid thinking with effort levels","detail":"One model id serves thinking (default) and non-thinking modes; reasoning effort low/high/max.","first":false,"discovered":"launch","source":"https://api-docs.deepseek.com/updates"},{"name":"Multiple API protocols","detail":"Same model served via OpenAI Chat Completions, OpenAI Responses (Codex-adapted) and Anthropic Messages formats.","first":false,"discovered":"launch","source":"https://api-docs.deepseek.com/quick_start/pricing"}],"entry":null,"notes":"Call as deepseek-flash. Legacy ids deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at Flash price. Knowledge cutoff not published.","verified":"2026-09-29","body_md":"DeepSeek's cheapest current model: agentic coding, long-context (1M) work and image understanding at very low cost. Schedule batch jobs off-peak for 50% off.\n\n```bash\ncurl https://api.deepseek.com/chat/completions \\\n -H \"Authorization: Bearer $DEEPSEEK_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"deepseek-flash\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://api-docs.deepseek.com/quick_start/pricing · https://api-docs.deepseek.com/updates · https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash","page_url":"https://postcutoff.com/m/deepseek-v4-1-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":448,"you_url":"https://postcutoff.com/you/deepseek-v4-1-flash/","briefings":null},{"id":"unifolm-wla-1-0","name":"UnifoLM-WLA-1.0","org":"Unitree Robotics","family":"UnifoLM","released":"2026-09-10","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"unitreerobotics/UnifoLM-WLA-1.0-Base","url":"https://huggingface.co/unitreerobotics/UnifoLM-WLA-1.0-Base"},{"provider":"Hugging Face (embodied reasoner backbones)","model_id":"unitreerobotics/UnifoLM-ER-Flow","url":"https://huggingface.co/unitreerobotics/UnifoLM-ER-1"},{"provider":"GitHub","url":"https://github.com/unitreerobotics/unifolm-wla","docs":"https://unigen-x.github.io/unifolm-wla.github.io/"}],"capabilities":[{"name":"One weight set for tabletop and whole-body humanoid manipulation","detail":"6B-parameter model coordinating 64 tasks across tabletop and whole-body manipulation on Unitree G1, with two-finger grippers and several five-finger dexterous hands.","first":false,"discovered":"launch","source":"https://github.com/unitreerobotics/unifolm-wla"},{"name":"Embodied reasoner + MMDiT action expert","detail":"Built on UnifoLM-ER (4B embodied reasoner based on Qwen3-VL-4B; 5M+ embodied reasoning samples) with an MMDiT action expert; ~2,500 h of real-robot data.","first":false,"discovered":"launch","source":"https://unigen-x.github.io/unifolm-wla.github.io/"}],"entry":"2026-09-10-unitree-unifolm-wla-1-0","notes":"Staged release: announcement + demo video 2026-09-10; UnifoLM-ER-1 / ER-Flow weights 2026-09-11; model modules and training code 2026-09-20; WLA-1.0-Base weights and fine-tuning code 2026-09-28 (GitHub news). HF repo lists Apache-2.0 but the model card was empty at check time. Predecessors: UnifoLM-VLA-0 (see unifolm-vla-0) and UnifoLM-WMA-0 world-model-action (Sept 2025). Benchmark claims (\"leading results across multiple embodied reasoning benchmarks\") are self-reported.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/unifolm-wla-1-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":448,"you_url":null,"briefings":null},{"id":"suno-v6","name":"Suno v6 (v6, v6-wild, v6-mini)","org":"Suno","family":"Suno v6","released":"2026-09-09","status":"current","type":"music","modality_in":["text","audio","image","video"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"free":0,"pro_per_month":10,"premier_per_month":30,"unit":"USD per month subscription (monthly billing; annual billing is 20% cheaper = $8 Pro / $24 Premier per month). Free: 50 credits/day, v6-mini only, no downloads, no commercial rights. Pro: 2,500 credits/month, 20 downloads/month, v6 + v6-wild, commercial rights. Premier: 10,000 credits/month, 60 downloads/month, Suno Studio.","source":"https://suno.com/pricing"},"price_line":"$10 per month, Pro plan","access":[{"provider":"Web app","url":"https://suno.com","model_id":"v6"},{"provider":"Web app (Pro/Premier)","model_id":"v6-wild"},{"provider":"Web app (all users, incl. free)","model_id":"v6-mini"}],"capabilities":[{"name":"Trained only on licensed music","detail":"First Suno generation developed with rightsholders; trained from scratch on music licensed from Warner Music Group, BMG and Believe (not on data used for earlier Suno versions), with revenue sharing to partners.","first":false,"discovered":"launch","source":"https://suno.com/blog/introducing-v6"},{"name":"Three-variant lineup","detail":"v6 (reliable, steerable flagship), v6-wild (experimental, genre-blending, pushes away from the prompt), v6-mini (fast, high-volume, available to everyone).","first":false,"discovered":"launch","source":"https://suno.com/blog/introducing-v6"},{"name":"Natural-language section and lyric editing","detail":"Edit parts of a song or change individual lyric lines by prompt without regenerating the whole track.","first":false,"discovered":"launch","source":"https://suno.com/blog/introducing-v6"},{"name":"Multimodal references and mashups","detail":"Text, audio, image and video references as a starting point; combine elements of several songs into a mashup; sample/isolate instruments and build beats.","first":false,"discovered":"launch","source":"https://suno.com/blog/introducing-v6"},{"name":"Upload screening and download limits","detail":"Uploaded audio and lyrics are screened for unauthorized use; downloads are capped per plan (none on Free, 20/month Pro, 60/month Premier).","first":false,"discovered":"launch","source":"https://suno.com/pricing"}],"entry":"2026-09-09-suno-v6-licensed-music-model","notes":"Launched 2026-09-09; Suno retired all earlier models (v4 to v5.5) as v6 rolled out. No official public API: in July 2026 Suno's CPO Jack Brody announced it was only 'exploring' a developer API/partner program (intake form, no timeline); third-party 'Suno APIs' are unofficial. Monthly-billing prices ($10/$30) are derived from the pricing page's annual price ($8/$24 per month) and its stated 20% annual discount. Max song length for v6 not stated on official pages checked. Sony Music and UMG sued again on 2026-09-18 over v6.","verified":"2026-09-29","body_md":"Suno's current song generator (vocals + full arrangement from a prompt/lyrics), the first Suno models trained on licensed music. Use via https://suno.com or the mobile apps; pick v6 / v6-wild (paid) or v6-mini (free) in the model selector.\n\nSources: [Introducing v6](https://suno.com/blog/introducing-v6), [pricing](https://suno.com/pricing), [TechCrunch](https://techcrunch.com/2026/09/09/suno-replaces-its-ai-models-with-a-new-one-trained-on-licensed-music-as-copyright-suits-pile-up/), [MBW on API exploration](https://www.musicbusinessworldwide.com/suno-explores-developer-api-seeking-apps-that-unlock-experiences-generative-music-makes-possible-for-the-first-time/).","page_url":"https://postcutoff.com/m/suno-v6/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":457,"you_url":null,"briefings":null},{"id":"yue2","name":"YuE2-3B","org":"Multimodal Art Projection (M-A-P)","family":"YuE","released":"2026-09-09","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"cc-by-nc-4.0 (weights; commercial license available), apache-2.0 (code)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/m-a-p/YuE2-3B"},{"provider":"GitHub (inference code, agent skill)","url":"https://github.com/multimodal-art-projection/YuE"},{"provider":"Hugging Face (community GGUF)","url":"https://huggingface.co/audio-cpp/Yue2-3B-GGUF"}],"capabilities":[{"name":"Score-first song generation","detail":"Writes an editable melody-and-chord plan in ABC notation, then renders a full song with vocals and accompaniment (48 kHz stereo).","first":false,"discovered":"launch","source":"https://github.com/multimodal-art-projection/YuE"},{"name":"Zero-shot covers and agentic editing","detail":"Covers from reference recordings (0.647 CLEWS mAP, self-reported) and conversational editing that turns musical feedback into score revisions.","first":false,"discovered":"launch","source":"https://huggingface.co/m-a-p/YuE2-3B"}],"entry":"2026-09-09-yue2-open-music-model","notes":"Self-reported WildSongBench best-of-8 6.9632 vs Suno v5 6.8721. English + Mandarin lyrics. Model card states ~4B parameters. HF repos created 2026-09-09; exact public announcement day not verified. Predecessor YuE (2025-01-28, arXiv 2503.08638).","verified":"2026-09-29","body_md":"Open-weights full-song generator from M-A-P (HKUST-led). Runs locally on one consumer GPU; weights download automatically on first use via the GitHub package.\n\nSources: https://github.com/multimodal-art-projection/YuE , https://huggingface.co/m-a-p/YuE2-3B","page_url":"https://postcutoff.com/m/yue2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":457,"you_url":null,"briefings":null},{"id":"gpt-image-2-5-flare","name":"GPT Image 2.5 Flare","org":"OpenAI","family":"GPT Image","released":"2026-09-08","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"text_input":5,"text_cached_input":1.25,"image_input":8,"image_cached_input":2,"image_output":30,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$5 text in, $30 image out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-image-2.5-flare","endpoint":"https://api.openai.com/v1/images/generations","docs":"https://developers.openai.com/api/docs/models/gpt-image-2.5-flare"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-image-2.5-flare","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"Web app","url":"https://chatgpt.com"},{"provider":"ElevenLabs Image & Video API","model_id":"gpt-image-2.5-flare","endpoint":"/image/create","docs":"https://elevenlabs.io/docs/overview/capabilities/image-video"}],"capabilities":[{"name":"Fast everyday image generation","detail":"Fastest high-quality OpenAI image model; quality low/medium/high/xhigh/max/auto.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-image-2.5-flare"},{"name":"Inpainting","detail":"Editing with masks via v1/images/edits.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-image-2.5-flare"}],"entry":"2026-09-08-chatgpt-images-2-5","notes":"Snapshot gpt-image-2.5-flare-2026-09-08. Same token rates as Sunburst and gpt-image-2. OpenRouter id not verified.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/images/generations \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-image-2.5-flare\",\"prompt\":\"a lighthouse at dawn\",\"quality\":\"medium\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-image-2.5-flare\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-image-2-5-flare/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":465,"you_url":null,"briefings":null},{"id":"gpt-image-2-5-sunburst","name":"GPT Image 2.5 Sunburst","org":"OpenAI","family":"GPT Image","released":"2026-09-08","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"text_input":5,"text_cached_input":1.25,"image_input":8,"image_cached_input":2,"image_output":30,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$5 text in, $30 image out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-image-2.5-sunburst","endpoint":"https://api.openai.com/v1/images/generations","docs":"https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-image-2.5-sunburst","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"Web app","url":"https://chatgpt.com"},{"provider":"ElevenLabs Image & Video API","model_id":"gpt-image-2.5-sunburst","endpoint":"/image/create","docs":"https://elevenlabs.io/docs/overview/capabilities/image-video"}],"capabilities":[{"name":"Most capable OpenAI image model","detail":"Top-quality generation and editing with inpainting via images/generations and images/edits.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst"},{"name":"Replacement for gpt-image-1.5/1-mini","detail":"Named successor for image models shutting down Dec 1 2026.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":"2026-09-08-chatgpt-images-2-5","notes":"Snapshot gpt-image-2.5-sunburst-2026-09-08. OpenRouter id not verified.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/images/generations \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-image-2.5-sunburst\",\"prompt\":\"a lighthouse at dawn\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-image-2-5-sunburst/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":465,"you_url":null,"briefings":null},{"id":"mercury-2-5","name":"Mercury 2.5","org":"Inception","family":"Mercury","released":"2026-09-08","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":260000,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":0.2,"output":0.75,"unit":"per 1M tokens (list price; 80% launch discount brings it to $0.04 / $0.15, end date not stated)","source":"https://docs.inceptionlabs.ai/get-started/models"},"price_line":"$0.20 in, $0.75 out per 1M tokens","access":[{"provider":"Inception API (OpenAI-compatible)","model_id":"mercury-2.5","endpoint":"https://api.inceptionlabs.ai/v1/chat/completions","docs":"https://docs.inceptionlabs.ai/get-started/models"},{"provider":"OpenRouter","model_id":"inception/mercury-2.5","url":"https://openrouter.ai/inception/mercury-2.5"},{"provider":"Baseten","url":"https://www.baseten.co/"}],"capabilities":[{"name":"Diffusion-based text generation at very high speed","detail":"A diffusion LLM (dLLM) that refines many tokens in parallel rather than one at a time; Inception claims 1,107 tokens/s on NVIDIA GPUs, and Artificial Analysis measured ~661 tokens/s.","first":false,"discovered":"launch","source":"https://www.inceptionlabs.ai/blog/introducing-mercury-2-5"},{"name":"Fast reasoning with tool calling","detail":"Configurable reasoning effort, tool calling and structured outputs; Inception claims a 40% intelligence gain over Mercury 2 and places it near GPT-5.6 Luna (low), Gemini 3.5 Flash-Lite and Claude Haiku 4.5.","first":false,"discovered":"launch","source":"https://www.inceptionlabs.ai/blog/introducing-mercury-2-5"}],"entry":null,"notes":"Artificial Analysis lists the release as Sept 8, 2026, measures ~661 tok/s and gives an Intelligence Index of 12, below average for its price tier. It lists $0.25/$0.75, while Inception's docs list $0.20/$0.75 before the discount. Inception's blog page showed Sept 29, 2026 when fetched, and OpenRouter also has a separate `inception/mercury-2.5-preview` listing, so the Sept 8 date may be the preview. Speed figures are vendor claims without disclosed batch or hardware details (RuntimeWire). Customer claims: Augment Code context compaction 150 s → 27 s; OpenCall ~170 ms median latency. 100M free API tokens for new accounts.","verified":"2026-09-30","body_md":"Mercury 2.5 is aimed at latency-sensitive work (voice agents, autocomplete, agent sub-calls) where speed matters more than peak intelligence.\n\nSources: [Inception blog](https://www.inceptionlabs.ai/blog/introducing-mercury-2-5) ·\n[Artificial Analysis](https://artificialanalysis.ai/models/mercury-2-5) ·\n[RuntimeWire](https://runtimewire.com/article/inception-mercury-2-5-diffusion-language-model-launch) ·\n[Inception on X](https://x.com/_inception_ai/status/2097365772289151417)","page_url":"https://postcutoff.com/m/mercury-2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":465,"you_url":"https://postcutoff.com/you/mercury-2-5/","briefings":null},{"id":"gpt-6-astra","name":"GPT-6 Astra","org":"OpenAI","family":"GPT-6","released":"2026-09-03","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-04","pricing":{"input":10,"cached_input":1,"output":50,"cache_write":12.5,"unit_note":"prompts >272K input tokens billed at 2x input and cache rates","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$10 in, $50 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-6-astra","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-6-astra"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-6-astra","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-6-astra","url":"https://openrouter.ai/openai/gpt-6-astra"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Max reasoning effort","detail":"reasoning.effort adds a new \"max\" level above xhigh (low/medium/high/xhigh/max).","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6-astra"},{"name":"1M-token context","detail":"1.05M context window (922K max input) with 128K output on the flagship.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-6-astra"},{"name":"Restricted cyber behaviour","detail":"Released as a restricted version that rejects certain cybersecurity prompts; separate Cyber/Daybreak models exist for that domain.","first":false,"discovered":"launch","source":"https://en.wikipedia.org/wiki/GPT-6_Astra"},{"name":"Recurrent-depth reasoning","detail":"Reported new \"recurrent depth\" technique that obscures some of the reasoning, raising monitorability concerns among safety researchers.","first":false,"discovered":"later","source":"https://en.wikipedia.org/wiki/GPT-6_Astra"},{"name":"Ultrafast service tier","detail":"From Sept 29 2026, service_tier \"ultrafast\" gives up to 6x faster generation in the API (8x / ~300 tok/s in Codex) at 6x price: $60 input / $6 cached / $75 cache write / $300 output per 1M tokens (<=272K ctx); default limits 500K-5M TPM.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/guides/ultrafast-mode"},{"name":"Powers dots always-on agents","detail":"OpenAI's dots (launched Sept 29 2026) run on GPT-6 Astra, each with its own cloud computer; also the default model in Agents API computer-use examples.","first":false,"discovered":"later","source":"https://openai.com/index/introducing-dots/"}],"entry":null,"notes":"OpenAI flagship (\"most capable model, built for the hardest end-to-end work\"). API changelog: Sep 3 2026 (limited preview Sep 3, public Sep 4). Single snapshot gpt-6-astra. OpenRouter also lists openai/gpt-6-astra-pro = same model with reasoning.mode pro. Endpoints: Chat Completions, Responses, Batch.","verified":"2026-09-29","body_md":"Hardest end-to-end work: long-horizon agentic coding, research, analysis, computer use.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-6-astra\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-6-astra\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://en.wikipedia.org/wiki/GPT-6_Astra\n- Ultrafast: https://developers.openai.com/api/docs/guides/ultrafast-mode\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-6-astra/","events_after":725,"major_after":180,"historic_after":38,"missing_at_launch":226,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-6-astra/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-04-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-04-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-04-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-04.md"}},{"id":"mai-transcribe-2","name":"MAI-Transcribe-2","org":"Microsoft","family":"MAI-Transcribe","released":"2026-09-03","status":"preview","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.1,"unit":"USD per hour of audio (limited-time promotional price through end of 2026; MAI-Transcribe-1.5 was $0.36/hr)","source":"https://microsoft.ai/models/mai-transcribe-2/"},"price_line":"$0.10 per hour","access":[{"provider":"Azure Speech in Microsoft Foundry (Fast Transcription API, enhancedMode)","model_id":"MAI-Transcribe-2","endpoint":"https://{resource}.cognitiveservices.azure.com/speechtotext/transcriptions:transcribe?api-version=2025-10-15","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe"},{"provider":"Azure Speech (previous version)","model_id":"MAI-Transcribe-1.5","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe"},{"provider":"Azure Voice Live (input transcription)","model_id":"MAI-Transcribe-2","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe"},{"provider":"OpenRouter","model_id":"microsoft/mai-transcribe-2","endpoint":"https://openrouter.ai/api/v1/audio/transcriptions"},{"provider":"Web app (MAI Playground)","url":"https://playground.microsoft.ai/"}],"capabilities":[{"name":"#1 on FLEURS across 60 languages (claimed)","detail":"Microsoft reports 5.2% average WER over 60 FLEURS languages (3.4% on top-25) and #2 on the Artificial Analysis WER leaderboard.","first":false,"discovered":"launch","source":"https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/"},{"name":"Very fast batch transcription","detail":"Claims ~10x faster than GPT-Transcribe (1 hour of audio in ~10 s), 7x vs Scribe v2, 5x vs Gemini 3.5.","first":false,"discovered":"launch","source":"https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/"},{"name":"Diarization, word timestamps, keyword biasing, clean/verbatim styles","detail":"New in v2: speaker diarization, word-level timestamps, phrase-list biasing, code-switching (e.g. Hinglish) and verbatim vs clean transcripts.","first":false,"discovered":"launch","source":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe"}],"entry":"2026-09-03-mai-transcribe-2","notes":"Public preview in Azure Speech. MAI-Transcribe-1.5 (Build 2026-06-02, 43 languages, $0.36/hr) remains available; MAI-Transcribe-1 deprecated 2026-08-20. Standard (post-promo) price not published. Input WAV/MP3/FLAC. Model card: https://microsoft.ai/pdf/MAI-Transcribe-2-Model-Card.pdf. Benchmarks are Microsoft-reported.","verified":"2026-09-29","body_md":"```bash\ncurl 'https://YOUR_RESOURCE.cognitiveservices.azure.com/speechtotext/transcriptions:transcribe?api-version=2025-10-15' \\\n  -H \"Ocp-Apim-Subscription-Key: $SPEECH_KEY\" -F 'audio=@meeting.wav' \\\n  -F 'definition={\"enhancedMode\":{\"enabled\":true,\"model\":\"MAI-Transcribe-2\"},\"diarization\":{\"enabled\":true}}'\n```\n\nSources: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-transcribe · https://microsoft.ai/models/mai-transcribe-2/ · https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/","page_url":"https://postcutoff.com/m/mai-transcribe-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":483,"you_url":null,"briefings":null},{"id":"muse-voice-transcribe-1-0","name":"Muse Voice Transcribe 1.0","org":"Meta","family":"Muse","released":"2026-09-03","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_minutes":3,"unit":"USD per 1,000 minutes of audio ($0.18/hour)","source":"https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/"},"price_line":"$3 per 1,000 minutes","access":[{"provider":"Meta Model API (streaming)","model_id":"muse-voice-transcribe-1.0","endpoint":"wss://api.meta.ai/v1/asr/realtime","docs":"https://dev.meta.ai/docs/overview"},{"provider":"Meta Model API (file)","model_id":"muse-voice-transcribe-1.0","endpoint":"https://api.meta.ai/v1/asr/transcribe","docs":"https://dev.meta.ai/docs/overview"}],"capabilities":[{"name":"#1 streaming STT on Artificial Analysis (claimed)","detail":"Meta says it ranks first on the Artificial Analysis streaming speech-to-text leaderboard and had the lowest average diarization error rate among APIs tested, streaming and offline.","first":false,"discovered":"launch","source":"https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/"},{"name":"Diarization, VAD and endpointing in one model","detail":"Speaker attribution for 20+ speakers, punctuation, speech-boundary detection and adaptive delay (uses more audio context only for ambiguous words).","first":false,"discovered":"launch","source":"https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/"}],"entry":"2026-09-03-meta-muse-voice-transcribe","notes":"Meta's first real-time audio perception model on the Meta Model API (launched 2026-09-03); 25+ languages. Speech-to-text only: Meta does not offer a TTS or speech-to-speech API; Muse's realtime voice mode and Muse Realtime Avatar (Connect, 2026-09-23) are consumer features without a documented API.","verified":"2026-09-29","body_md":"Streaming and file transcription for voice agents built on Muse.\n\nSources: https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/ · https://dev.meta.ai/docs/overview","page_url":"https://postcutoff.com/m/muse-voice-transcribe-1-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":483,"you_url":null,"briefings":null},{"id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","org":"Google DeepMind","family":"Gemini 3","released":"2026-09-02","status":"current","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":"2026-03","pricing":{"input":0.75,"output":3.75,"unit":"per 1M tokens (Standard tier, introductory price through 2026-12-31; rises to $1.50 / $7.50 from 2027-01-01)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.75 in, $3.75 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.8-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"gemini-3.8-flash","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-8-flash"},{"provider":"OpenRouter","model_id":"google/gemini-3.8-flash"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Long-horizon software engineering","detail":"Google's most capable Flash for autonomous end-to-end engineering; 73.7% on DeepSWE v1.1, 89.4% terminal-based coding per model card.","first":false,"discovered":"launch","source":"https://deepmind.google/models/model-cards/gemini-3-8-flash/"},{"name":"Specialized-domain agentic analysis","detail":"Beats 3.7 Flash and other frontier models on Vals Finance Agent v2 (61.4%) and Harvey's Legal Agent Benchmark.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"},{"name":"Agentic long-video understanding","detail":"87.8% long video understanding in agentic mode (agentic video understanding added for 3.x Flash on 2026-09-01).","first":false,"discovered":"launch","source":"https://deepmind.google/models/model-cards/gemini-3-8-flash/"},{"name":"Adjustable thinking levels + computer use","detail":"Thinking low/medium/high, computer use (preview), Maps/Search grounding, flex and priority inference tiers.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash"},{"name":"Cyber sibling model","detail":"Launched alongside Gemini 3.8 Flash Cyber (vulnerability detection/patching, 47.2% CWE-Bench pass@1), available only to vetted defenders via the Fairwind Program.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"}],"entry":null,"notes":"Newest and recommended Gemini text model as of 2026-09 (no Pro newer than 3.1 Pro preview; 3.5 Pro announced but unreleased). Aliases gemini-flash-latest may point here. Model card says some domains' knowledge only to 2025-01. Vertex id inferred from docs page.","verified":"2026-09-29","body_md":"Google's flagship workhorse model (Sept 2026): best Gemini for coding agents, multi-step reasoning and long multimodal context at Flash prices.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Explain how AI works in a few words\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/), [model card](https://deepmind.google/models/model-cards/gemini-3-8-flash/).","page_url":"https://postcutoff.com/m/gemini-3-8-flash/","events_after":747,"major_after":187,"historic_after":40,"missing_at_launch":243,"events_since_release":null,"you_url":"https://postcutoff.com/you/gemini-3-8-flash/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-03-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-03-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-03-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-03.md"}},{"id":"muse-spark-1-3","name":"Muse Spark 1.3","org":"Meta","family":"Muse","released":"2026-09-02","status":"current","type":"reasoning-llm","modality_in":["text","image","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":200000,"knowledge_cutoff":null,"pricing":{"input":1.25,"cached_input":0.15,"output":4.25,"unit":"per 1M tokens (USD), standard tier; \"contributor\" tier muse-spark-1.3-contributor is $0.10/$0.002 cached/$0.20","source":"https://dev.meta.ai/models/muse-spark/"},"price_line":"$1.25 in, $4.25 out per 1M tokens","access":[{"provider":"Meta Model API","model_id":"muse-spark-1.3","endpoint":"https://api.meta.ai/v1/chat/completions","docs":"https://dev.meta.ai/docs/"},{"provider":"OpenRouter","model_id":"meta/muse-spark-1.3","url":"https://openrouter.ai/meta/muse-spark-1.3"},{"provider":"Web app","url":"https://meta.ai"}],"capabilities":[{"name":"Closed-weights successor to Llama","detail":"Proprietary model from Meta Superintelligence Labs; Muse Spark replaced Llama in Meta AI in April 2026.","first":false,"discovered":"launch","source":"https://venturebeat.com/technology/goodbye-llama-meta-launches-new-proprietary-ai-model-muse-spark-first-since"},{"name":"Native video + document perception","detail":"Natively multimodal input (video, images, documents, text) with 1M context and 200K max output.","first":false,"discovered":"launch","source":"https://dev.meta.ai/models/muse-spark/"},{"name":"Long-horizon multi-agent tuning","detail":"1.3 tuned for long-running, multi-agent agentic builds; also powers Meta's Muse Code.","first":false,"discovered":"launch","source":"https://x.com/MetaforDevs/status/2095232442953236714"},{"name":"Contributor pricing tier","detail":"Separate -contributor model ids priced ~90% lower (data-sharing tier).","first":false,"discovered":"launch","source":"https://dev.meta.ai/docs/"}],"entry":"2026-09-02-meta-muse-spark-1-3","notes":"Other ids: muse-spark-1.2, muse-spark-1.1, muse-spark-1.3-contributor, muse-spark-1.2-contributor. OpenAI-SDK-compatible API (public preview, self-serve). Contributor tier data terms not verified. Knowledge cutoff not published.","verified":"2026-09-29","body_md":"Meta's current flagship (closed) for agentic coding and multimodal (video/doc) understanding.\n\n```bash\ncurl https://api.meta.ai/v1/chat/completions -H \"Authorization: Bearer $MODEL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"muse-spark-1.3\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://dev.meta.ai/models/muse-spark/ , https://dev.meta.ai/docs/ , https://research.meta.ai/blog/introducing-muse-spark-1-3","page_url":"https://postcutoff.com/m/muse-spark-1-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":499,"you_url":"https://postcutoff.com/you/muse-spark-1-3/","briefings":null},{"id":"claude-fable-5-1","name":"Claude Fable 5.1","org":"Anthropic","family":"Claude 5","released":"2026-09-01","status":"current","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-06","pricing":{"input":10,"output":50,"cache_read":0.25,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$10 in, $50 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-fable-5-1","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/fable-5-1/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-fable-5-1","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-fable-5-1","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-fable-5-1","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-fable-5-1","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-fable-5.1","url":"https://openrouter.ai/anthropic/claude-fable-5.1"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Scientific discovery (protein design)","detail":"In Anthropic's launch examples, its protein designs reached about 10x higher binding affinity than competition winners, with a hit rate near 50%.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"name":"Rare-bug hunting","detail":"Anthropic reports it found the cause of a one-in-a-million crash that engineers had not explained for years.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"name":"Top CursorBench score","detail":"Scored 73.4% on CursorBench 3.2.0 at max effort, which Cursor called the most capable model it had run.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"name":"Preserved thinking and content provenance","detail":"Thinking blocks are bound to the model and the conversation, and editing earlier turns invalidates them. Also adds per-message effort, turn-scoped system messages and content provenance.","first":true,"discovered":"launch","source":"https://platform.claude.com/docs/en/models/fable-5-1/overview"},{"name":"Cheaper cache reads","detail":"Cache reads cost $0.25/MTok (0.025x input). Anthropic cites up to 45% savings on agentic work compared with Fable 5.","first":false,"discovered":"launch","source":"https://platform.claude.com/docs/en/about-claude/pricing"}],"entry":null,"notes":"Anthropic's most capable widely released model; thinking always on (adaptive, effort low..max, default high); forced tool_choice any/tool returns 400; no prefill; 30-day data retention required (no ZDR unless authorized); no Priority Tier. Batch $5/$25.","verified":"2026-09-29","body_md":"Best for the hardest reasoning and long-horizon agentic work (coding, research, documents/spreadsheets/slides, computer use). Successor to Claude Fable 5 at the same price with cheaper cache reads. Same model is offered as Claude Mythos 5.1 to Project Glasswing participants. Handle `stop_reason: \"refusal\"` (safety classifiers) and consider the server-side `fallbacks` parameter.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-fable-5-1\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/fable-5-1/overview\n- Announcement: https://www.anthropic.com/claude-fable-and-mythos-5-1\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-fable-5-1/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":157,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-fable-5-1/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"claude-mythos-5-1","name":"Claude Mythos 5.1","org":"Anthropic","family":"Claude 5","released":"2026-09-01","status":"current","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-06","pricing":{"input":10,"output":50,"cache_read":0.25,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$10 in, $50 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-mythos-5-1","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/mythos-5-1/overview"}],"capabilities":[{"name":"Frontier cyber-defense model","detail":"Offered only to Project Glasswing participants for defensive cybersecurity. It has the same capabilities as Fable 5.1, with safeguards that depend on the access program.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"name":"Scientific discovery","detail":"Shares Fable 5.1's launch results, e.g. protein designs with about 10x higher binding affinity than competition winners.","first":false,"discovered":"launch","source":"https://www.anthropic.com/claude-fable-and-mythos-5-1"}],"entry":null,"notes":"Invitation-only (Project Glasswing, defensive cybersecurity). Same capabilities/pricing as Claude Fable 5.1; not offered on Claude Platform on AWS. Cloud ids not listed publicly; contact Anthropic/AWS/Google account team. Successor to claude-mythos-5 and claude-mythos-preview (deprecated 2026-06-09).","verified":"2026-09-29","body_md":"Same model as Claude Fable 5.1 offered to approved Project Glasswing participants (https://anthropic.com/glasswing). Unlike Fable 5.1 it does not run the preserved-thinking history-editing check. If your org is not in Glasswing, use `claude-fable-5-1`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-mythos-5-1\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/mythos-5-1/overview\n- Announcement: https://www.anthropic.com/claude-fable-and-mythos-5-1\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-mythos-5-1/","events_after":665,"major_after":165,"historic_after":36,"missing_at_launch":157,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-mythos-5-1/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-06.md"}},{"id":"qwen3-8-omni-flash","name":"Qwen3.8-Omni-Flash","org":"Alibaba (Qwen)","family":"Qwen3.8","released":"2026-09","status":"current","type":"multimodal","modality_in":["text","image","audio","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":131072,"knowledge_cutoff":null,"pricing":{"input":0.15,"output":0.47,"cache_read":0.016,"unit":"per 1M tokens (USD), Singapore/International","source":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},"price_line":"$0.15 in, $0.47 out per 1M tokens","access":[{"provider":"Alibaba Cloud Model Studio (DashScope, Singapore/Intl)","model_id":"qwen3.8-omni-flash","endpoint":"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash"},{"provider":"Alibaba Cloud Model Studio (realtime voice/video)","model_id":"qwen3.8-omni-flash-realtime","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"OpenRouter","model_id":"qwen/qwen3.8-omni-flash","url":"https://openrouter.ai/qwen/qwen3.8-omni-flash"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"Audio + video understanding with 1M context","detail":"Text, image, audio and video in, text out; 113 input languages/dialects for audio.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash"},{"name":"Spatial (multichannel) audio input","detail":"Accepts multichannel/spatial audio via use_multichannel in Chat Completions.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash"},{"name":"Realtime speech-to-speech sibling","detail":"qwen3.8-omni-flash-realtime handles live audio/video conversation; for non-realtime audio output Alibaba points to qwen3.5-omni-plus.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/models"}],"entry":"2026-09-18-qwen3-8-omni-flash","notes":"Thinking on by default with adjustable effort. Realtime variant qwen3.8-omni-flash-realtime: $0.93 audio in / $1.87 audio out per 1M tokens (Singapore/Intl pricing page, checked 2026-09-29). For dedicated hosted voice agents Alibaba also offers qwen-audio-3.1-realtime-plus (see qwen-audio-3-1-realtime). Official blog: https://qwen.ai/blog?id=qwen3.8-omni-flash (Qwen blog index date 2026-09-18; the page header shows Sept 14 as a draft); OpenRouter listing 2026-09-21.","verified":"2026-09-29","body_md":"Alibaba's omni model for transcription-plus-reasoning, meeting/video analysis and multimodal agents at Flash prices.\n\n```bash\ncurl \"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions\" \\\n -H \"Authorization: Bearer $DASHSCOPE_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"qwen3.8-omni-flash\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nFor audio/video inputs see the non-real-time guide linked from the model page.\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-omni-flash · https://www.alibabacloud.com/help/en/model-studio/model-pricing","page_url":"https://postcutoff.com/m/qwen3-8-omni-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":165,"you_url":"https://postcutoff.com/you/qwen3-8-omni-flash/","briefings":null},{"id":"inworld-tts-2","name":"Inworld Realtime TTS-2 / TTS-2 Flash","org":"Inworld AI","family":"Inworld TTS","released":"2026-08-31","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":25,"unit":"USD per 1M characters PAYG for TTS-2 ($15 for TTS-2 Flash); plan rates down to $12.50/$7, enterprise from $5","source":"https://inworld.ai/pricing"},"price_line":"$25 per 1M characters","access":[{"provider":"Inworld API","model_id":"inworld-tts-2","endpoint":"https://api.inworld.ai/tts/v1/voice","docs":"https://docs.inworld.ai/tts/tts-models"},{"provider":"Cloudflare Workers AI","url":"https://developers.cloudflare.com/ai/models/inworld/tts-2/"}],"capabilities":[{"name":"Closed-loop, audio-aware delivery","detail":"Conditions on the actual audio of prior turns (user tone, pacing, emotion), not just transcripts, and takes plain-English voice direction; delivery modes STABLE/BALANCED/CREATIVE.","first":false,"discovered":"launch","source":"https://inworld.ai/blog/realtime-tts-2"},{"name":"Cross-lingual identity in 100+ languages","detail":"One voice holds identity while switching language on the fly; cloning from 5-15 s reference or voice design from a text description.","first":false,"discovered":"launch","source":"https://inworld.ai/blog/realtime-tts-2"},{"name":"Flash variant ~20 ms TTFB","detail":"TTS-2 Flash: ~20 ms TTFB, ~5x faster than inworld-tts-2 (docs); TTS-2 median TTFA under 200 ms.","first":false,"discovered":"launch","source":"https://docs.inworld.ai/tts/tts-models"}],"entry":"2026-08-31-inworld-realtime-tts-2","notes":"Research preview 2026-05-05, GA 2026-08-31. Inworld claimed #1 on Artificial Analysis Speech Arena; on 2026-09-29 AA shows it #5 (Elo 1244) behind Eleven v4, Sonic 3.6, Gemini 3.8 Flash TTS, Qwen-Audio-3.0-TTS-Plus. Docs say 200+ languages vs 100+ in blog. TTS-1..1.5 discontinued 2026-06-15 (auto-routed). Flash model id not verified. Max 2,000 chars/request.","verified":"2026-09-29","body_md":"Sources: https://inworld.ai/blog/realtime-tts-2 , https://docs.inworld.ai/tts/tts-models , https://inworld.ai/pricing","page_url":"https://postcutoff.com/m/inworld-tts-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":508,"you_url":null,"briefings":null},{"id":"cartesia-sonic-3-6","name":"Cartesia Sonic-3.6","org":"Cartesia","family":"Sonic","released":"2026-08-27","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.037,"unit":"approx., derived: ~1 credit per character; Scale plan $299/mo for 8M credits (Startup $49 for 1.25M ≈ $0.039/1K). Artificial Analysis lists $49 per 1M chars","source":"https://cartesia.ai/pricing"},"price_line":"$0.037 per 1K characters","access":[{"provider":"Cartesia API","model_id":"sonic-3.6","endpoint":"https://api.cartesia.ai/tts/bytes (also /tts/sse and WebSocket)","docs":"https://docs.cartesia.ai/build-with-cartesia/tts-models/latest"},{"provider":"Cartesia API (pinned snapshot)","model_id":"sonic-3.6-2026-08-27","endpoint":"https://api.cartesia.ai/tts/bytes"},{"provider":"Web app","url":"https://play.cartesia.ai"}],"capabilities":[{"name":"State-space-model TTS, sub-90 ms","detail":"Built on state space models (SSMs); replies in under 90 ms and generates ~132 chars/s (nearly 2x Sonic 3 Conversational). Listeners preferred it over Sonic-3.5 in up to 93% of blind tests across 15 locales.","first":false,"discovered":"launch","source":"https://www.cartesia.ai/blog/sonic-3.6"},{"name":"44 languages with instant cloning","detail":"Adds Odia and Urdu to Sonic-3.5's 42 languages; instant voice cloning; locale-aware reading of dates/numbers; confirmation codes and heteronyms without preprocessing.","first":false,"discovered":"launch","source":"https://docs.cartesia.ai/build-with-cartesia/tts-models/latest"},{"name":"Multilingual Voices (one voice, 25 languages)","detail":"Launched 2026-09-23 on Sonic-3.6: 50+ library voices each speak up to 25 languages natively, and custom clones from ~10 s of audio carry their identity across languages via a `locale` parameter; native speakers rate each variant for accent and localization of dates, numbers and currency.","first":false,"discovered":"later","source":"https://www.cartesia.ai/blog/multilingual-voices"},{"name":"Top-2 on Artificial Analysis Speech Arena","detail":"Ranked #1 (Elo ~1279) on the Artificial Analysis TTS leaderboard in mid/late Sept 2026, then #2 (Elo 1275) behind Eleven v4 after 2026-09-28.","first":false,"discovered":"later","source":"https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"}],"entry":"2026-08-27-cartesia-sonic-3-6","notes":"Beta 2026-08-17, GA snapshot 2026-08-27. Header `Cartesia-Version: 2026-08-14`. Fully backwards compatible with Sonic-3.5 (snapshot 2026-05-04, which led AA's Controlled Voice Arena at its 2026-07-08 launch with 1,122 Elo). Scored 0.840 (#5) on Hume's Real-World VoiceEQ leaderboard (2026-09-24). sonic-3 snapshots (2025-10-27, 2026-01-12), sonic-2 and sonic-turbo sunset 2026-10-20. `sonic-preview` = beta channel; `sonic-latest` alias deprecated. Exact per-character USD price is plan-dependent (credits); figure above is derived. Also on AWS SageMaker JumpStart (Sonic 3, Feb 2026).","verified":"2026-09-29","body_md":"Cartesia's flagship low-latency TTS for voice agents.\n\n```bash\ncurl -X POST https://api.cartesia.ai/tts/bytes -H \"Authorization: Bearer $CARTESIA_API_KEY\" -H \"Cartesia-Version: 2026-08-14\" \\\n  -H \"Content-Type: application/json\" -d '{\"model_id\":\"sonic-3.6\",\"transcript\":\"Your code is 4 8 1 5.\",\"voice\":{\"mode\":\"id\",\"id\":\"<voice_id>\"},\"output_format\":{\"container\":\"wav\",\"encoding\":\"pcm_s16le\",\"sample_rate\":44100}}' -o out.wav\n```\n\nSources: https://docs.cartesia.ai/build-with-cartesia/tts-models/latest , https://docs.cartesia.ai/build-with-cartesia/tts-models/older-models , https://www.cartesia.ai/blog/sonic-3.6","page_url":"https://postcutoff.com/m/cartesia-sonic-3-6/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":517,"you_url":null,"briefings":null},{"id":"gemini-3-5-transcribe","name":"Gemini 3.5 Transcribe (and Transcribe Live)","org":"Google DeepMind","family":"Gemini 3.5","released":"2026-08-26","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"audio_input":2,"output":12,"unit":"per 1M tokens (USD) for gemini-3.5-transcribe (~$0.003/min audio in + ~$0.002/min text out); gemini-3.5-transcribe-live $3.50 in / $21.00 out (~$0.005 + ~$0.004 per min)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$2 audio in, $12 out per 1M tokens","access":[{"provider":"Gemini API (Interactions API, files)","model_id":"gemini-3.5-transcribe","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe"},{"provider":"Gemini Live API (WebSocket streaming)","model_id":"gemini-3.5-transcribe-live","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe"},{"provider":"Google AI Studio","url":"https://aistudio.google.com"}],"capabilities":[{"name":"Smart transcription","detail":"Handles self-corrections, removes filler words and auto-formats text; custom vocabulary biasing up to 1,000 terms.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/"},{"name":"Low word error rate","detail":"Google cites Artificial Analysis WER of 2.6% (non-streaming) and 4.0% (streaming); 70% faster time-to-final than Chirp 3.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/"},{"name":"85+ languages with code-switching, diarization, word timestamps","detail":"Utterance-level language detection across 85+ languages; speaker diarization; word-level timestamps (not combinable with custom vocabulary).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe"}],"entry":null,"notes":"Changelog lists both ids GA on 2026-08-26, while the launch blog says public preview in AI Studio and Gemini Enterprise Agent Platform. Limits: 1 h per file request (30 min with diarization/timestamps), 10 min per live session. Diarization: docs say up to 8 speakers, blog says up to three - unresolved. Powers Rambler on Android and the Gemini app on macOS; coming to Chrome and Gboard. Press quotes $0.005/min (file) and $0.009/min (live) all-in. Model card (read 2026-09-29): https://deepmind.google/models/model-cards/gemini-3-5-audio/ - covers Gemini 3.5 Live Translate, Transcribe and Transcribe Live (card dated 2026-08-26); no numeric evals in the card itself; knowledge cutoff January 2025; did not reach any Tracked or Critical Capability Levels under the Frontier Safety Framework. Card lists hallucinations and occasional slowness/timeouts as limitations; surfaces: Antigravity, Gboard, Gemini app, Vertex AI, Google Workspace.","verified":"2026-09-29","body_md":"Google's Gemini-based speech-to-text, successor in practice to Cloud [Chirp 3](chirp-3.md) for developers.\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [changelog](https://ai.google.dev/gemini-api/docs/changelog), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/).","page_url":"https://postcutoff.com/m/gemini-3-5-transcribe/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":524,"you_url":null,"briefings":null},{"id":"breeze-tts-2","name":"Breeze TTS 2","org":"BreezeBlue","family":"Breeze TTS","released":"2026-08-25","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"BreezeBlue Research and Non-Commercial License (weights); Apache-2.0 (code)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"BreezeBlue/Breeze-TTS-2","url":"https://huggingface.co/BreezeBlue/Breeze-TTS-2"},{"provider":"GitHub","url":"https://github.com/breezeblue-ai/breeze-tts"},{"provider":"BreezeBlue (hosted / commercial license)","url":"https://breezeblue.ai"}],"capabilities":[{"name":"#1 open-weights TTS on Artificial Analysis","detail":"~1,206-1,215 Elo in the Artificial Analysis Speech Arena, ~90 points above Fish Audio S2 Pro, #6 overall at launch — the leading open-weights TTS as of Sept 2026.","first":false,"discovered":"later","source":"https://x.com/ArtificialAnlys/status/2092399623839326550"},{"name":"Clone + design + direct in one 3B checkpoint, <40 ms TTFA","detail":"Voice cloning from reference audio, voice design from text descriptions, voice direction (tone/emotion keeping identity), vocal events (laughs, coughs); streaming TTFA under 40 ms on H100 with fast path, RTF 0.32.","first":false,"discovered":"launch","source":"https://huggingface.co/BreezeBlue/Breeze-TTS-2"}],"entry":"2026-08-25-breeze-tts-2","notes":"Model card lists English + Chinese; the Artificial Analysis post mentions 50 languages (possibly the hosted model) — unresolved. Needs 12 GB VRAM (24 GB recommended), CUDA/Linux. Weights are NOT commercially usable without a BreezeBlue subscription. Some secondary blogs claim it is the 'first open-weight model to beat ElevenLabs' flagship' — unverified and contradicted by the AA leaderboard (Eleven v4 far ahead).","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/BreezeBlue/Breeze-TTS-2 , https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice/open-weights","page_url":"https://postcutoff.com/m/breeze-tts-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":532,"you_url":null,"briefings":null},{"id":"skild-s1","name":"Skild S1 (Skild Brain)","org":"Skild AI","family":"Skild Brain","released":"2026-08-25","status":"current","type":"robotics","modality_in":["video","image","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Skild AI (commercial partners; early-access sign-up)","url":"https://www.skild.ai/blogs/s1"}],"capabilities":[{"name":"In-context learning from one video, long-horizon","detail":"Learns tasks never seen in pretraining (potting a plant, cooking pancakes, pour-over coffee, kit assembly) from a single video prompt with no fine-tuning, for tasks up to ~10 minutes long; Skild calls this the first robotics foundation model to show in-context learning on such long unseen tasks.","first":true,"discovered":"launch","source":"https://www.skild.ai/blogs/s1"},{"name":"Video prompting beats language prompting","detail":"66% success on unseen tasks vs 9% for an equivalently trained language-prompted policy (~7x); 96% on seen tasks; one demo video worth ~380 post-training episodes; 11 minutes from demonstration to autonomous execution in the plant-potting example.","first":false,"discovered":"launch","source":"https://www.skild.ai/blogs/s1"},{"name":"Omni-bodied brain","detail":"Skild Brain is pitched as one model controlling quadrupeds, humanoids, arms and mobile manipulators without prior knowledge of the body; S1 trains on teleop, human video, simulation and data-capture gloves.","first":false,"discovered":"launch","source":"https://www.therobotreport.com/skild-ai-unveils-s1-flagship-robot-foundation-model/"}],"entry":"2026-08-25-skild-ai-s1","notes":"Announced on X 2026-08-25 (https://x.com/SkildAI/status/2092300842900865389); press 2026-08-31; NVIDIA blog 2026-09-10 (https://blogs.nvidia.com/blog/skild-ai-s1-physical-ai/) cites a $100M revenue run rate 10 months after first commercial deployment, 60+ deployment partnerships and Blackwell assembly work with Foxconn. Skild raised a $1.4B Series C at >$14B (2026-01-14, led by SoftBank). No public API, pricing or weights; company says S1 is \"already at work with our commercial partners\" and plans wider real-world rollout by 2027. Results are company-reported.","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/skild-s1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":532,"you_url":null,"briefings":null},{"id":"elevenlabs-v3-conversational","name":"Eleven v3 Conversational","org":"ElevenLabs","family":"Eleven v3","released":"2026-08-19","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.04,"unit":"per 1K characters (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.04 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_v3_conversational","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}/stream","docs":"https://elevenlabs.io/docs/models"},{"provider":"ElevenAgents","url":"https://elevenlabs.io/agents"}],"capabilities":[{"name":"Real-time v3 with audio tags","detail":"Brings Eleven v3's expressive delivery and audio tags to streaming/real-time use at ~280 ms latency (excl. application & network), 70+ languages.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"}],"entry":null,"notes":"GA announced 2026-08-19 (ElevenLabs X post and ElevenLabs Developers YouTube video). Artificial Analysis TTS arena Elo ~1196 (Aug 2026). Superseded for agents by eleven_v4_turbo (2026-09-28, ~100 ms). Exact streaming endpoint shown is the generic TTS stream endpoint; websockets also used in ElevenAgents.","verified":"2026-09-29","body_md":"Real-time variant of Eleven v3 for voice agents.\n\nSources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://x.com/ElevenLabs/status/2090136227617952145 , https://www.youtube.com/watch?v=pNMYYsO_UBE","page_url":"https://postcutoff.com/m/elevenlabs-v3-conversational/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":546,"you_url":null,"briefings":null},{"id":"generalist-gen-1-5","name":"Generalist GEN-1.5","org":"Generalist AI","family":"GEN","released":"2026-08-19","status":"current","type":"robotics","modality_in":["video","image","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Generalist AI partners (no public access announced)","url":"https://generalistai.com/blog/gen-1.5"}],"capabilities":[{"name":"One-shot learning of dexterous closed-loop tasks","detail":"Learns new tasks in-context from one demonstration video: 59% average success one-shot across 10 tasks; 83% with few-shot adaptation (10 gradient steps on 5 minutes of data). Generalist says it is the first model it knows of to show this across a wide range of dexterous closed-loop tasks.","first":true,"discovered":"launch","source":"https://generalistai.com/blog/gen-1.5"},{"name":"30-second video memory, 100 Hz actions","detail":"Takes video (30 s memory window), sensors, language and proprioception and outputs 100 Hz action trajectories.","first":false,"discovered":"launch","source":"https://generalistai.com/blog/gen-1.5"}],"entry":"2026-08-19-generalist-gen-1-5","notes":"Released 6 days before Skild S1, which makes a similar one-video in-context claim for long-horizon tasks. Company-reported. Video: https://www.youtube.com/watch?v=1cllCVK-9lo","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/generalist-gen-1-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":546,"you_url":null,"briefings":null},{"id":"teleocr","name":"TeleOCR (ex NaviDC-OCR)","org":"TeleAI","family":"TeleOCR","released":"2026-08-17","status":"current","type":"multimodal","modality_in":["image","pdf"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/XingChen-AGI/TeleOCR"},{"provider":"GitHub","url":"https://github.com/caipeng328/TeleOCR"},{"provider":"Web app","url":"https://www.teleai.com.cn/docparse/DocumentParsing"}],"capabilities":[{"name":"Camera-captured document parsing without dewarping","detail":"One ~1.4B-parameter Qwen2.5-VL-based model parses both digital and photographed or curved documents directly, with no separate rectification step (DocUNet/DIR300 demos).","first":false,"discovered":"launch","source":"https://huggingface.co/XingChen-AGI/TeleOCR"},{"name":"Dr.DocBench Challenge lead","detail":"Self-reported 67.96 overall on the EMNLP 2026 Dr.DocBench Challenge vs MinerU 2.5 Pro 62.26 and PaddleOCR-VL 1.6 55.11.","first":false,"discovered":"later","source":"https://huggingface.co/XingChen-AGI/TeleOCR"}],"entry":null,"notes":"Released Aug 17, 2026 as NaviDC-OCR (StarDoc-AI), renamed TeleOCR on Sept 10. ~1,280 HF likes and ~41k downloads by Oct 6, 2026; GGUF/llama.cpp ports exist. Benchmarks are self-reported. The README links TeleAI's (China Telecom) document-parsing web app; the HF org is XingChen-AGI.","verified":"2026-10-06","body_md":"Lightweight open document-parsing VLM. Technical report: arXiv 2608.12898.","page_url":"https://postcutoff.com/m/teleocr/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":554,"you_url":"https://postcutoff.com/you/teleocr/","briefings":null},{"id":"gemini-3-7-flash","name":"Gemini 3.7 Flash","org":"Google DeepMind","family":"Gemini 3","released":"2026-08-13","status":"legacy","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":0.75,"output":3.75,"unit":"per 1M tokens (Standard tier, introductory through 2026-12-31; $1.50 / $7.50 from 2027-01-01)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.75 in, $3.75 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.7-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-7-flash"},{"provider":"OpenRouter","model_id":"google/gemini-3.7-flash"}],"capabilities":[{"name":"Production-quality coding","detail":"43.6% FrontierCode 1.1 Main and 65.3% DeepSWE v1.1 (vs 34.4% / 49.0% for 3.6 Flash); WebDev Arena Elo 1588.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/"},{"name":"Enterprise document/automation work","detail":"34.0% GDP.pdf and 30.4% AutomationBench, large jumps over 3.6 Flash.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/"},{"name":"Half-price workhorse","detail":"Launched at half the original 3.6 Flash per-token price; thinking levels low/medium/high (minimal returns an error).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash"}],"entry":null,"notes":"Still served and stable (no shutdown date) but superseded by Gemini 3.8 Flash at the same price. Vertex model id not verified.","verified":"2026-09-29","body_md":"Previous-generation Flash (Aug 2026). Prefer `gemini-3.8-flash` for new work; same price.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/).","page_url":"https://postcutoff.com/m/gemini-3-7-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":561,"you_url":null,"briefings":null},{"id":"deepgram-flux-tts","name":"Deepgram Flux TTS","org":"Deepgram","family":"Flux","released":"2026-08-12","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.045,"unit":"per 1K characters pay-as-you-go ($0.0405 Growth); free until 2026-09-12, standard pricing from 2026-09-13","source":"https://deepgram.com/pricing"},"price_line":"$0.045 per 1K characters","access":[{"provider":"Deepgram API (real-time)","model_id":"flux-haley-en","endpoint":"wss://api.deepgram.com/v2/speak?model=flux-haley-en","docs":"https://developers.deepgram.com/docs/flux-tts/overview"},{"provider":"Deepgram API (batch)","model_id":"flux-{voice}-en","endpoint":"https://api.deepgram.com/v2/speak","docs":"https://developers.deepgram.com/docs/flux-tts/voices"}],"capabilities":[{"name":"Conversation-native TTS","detail":"Keeps context and voice consistency across turns of a conversation instead of treating each sentence in isolation; turn lifecycle events; on Interrupt reports exactly what the user heard (`text_spoken`).","first":false,"discovered":"launch","source":"https://deepgram.com/learn/text-to-speech-comes-of-age-deepgram-launches-conversation-native-speech"},{"name":"~80 ms response, structured-content accuracy","detail":"Starts responding in as little as 80 ms under production load; tuned for account numbers, alphanumerics, drug names and money amounts.","first":false,"discovered":"launch","source":"https://deepgram.com/learn/text-to-speech-comes-of-age-deepgram-launches-conversation-native-speech"}],"entry":"2026-08-12-deepgram-flux-tts","notes":"English only (39 voices; American, British, Irish, Australian, Indian, Singaporean, Filipino accents); use Aura-2 for other languages. Self-hosted GA 2026-08-26; speed 0.5-1.5 and expressivity -2..2 controls. Launched alongside Deepgram passing $100M ARR.","verified":"2026-09-29","body_md":"Voice-agent TTS served on /v2/speak (WebSocket for streaming LLM tokens, REST for batch).\n\nSources: https://developers.deepgram.com/docs/flux-tts/overview , https://developers.deepgram.com/docs/flux-tts/voices , https://deepgram.com/pricing , https://developers.deepgram.com/changelog","page_url":"https://postcutoff.com/m/deepgram-flux-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":566,"you_url":null,"briefings":null},{"id":"ltx-2-5","name":"LTX-2.5","org":"Lightricks","family":"LTX-2","released":"2026-08-11","status":"current","type":"video-gen","modality_in":["text","image","video"],"modality_out":["video","audio"],"open_weights":true,"model_license":"ltx-2-community-license (free under $10M annual revenue)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"fast_per_second":0.09,"4k_per_second":0.3,"unit":"USD per second of output video on Lightricks' managed API (as reported by VentureBeat)","source":"https://venturebeat.com/technology/ltx-2-5-can-generate-a-10-second-ai-video-from-an-image-in-just-6-8-seconds-on-nvidia-superchips-and-its-open-weights"},"price_line":"$0.09 per second, fast","access":[{"provider":"Hugging Face","url":"https://huggingface.co/Lightricks/LTX-2.5"},{"provider":"GitHub","url":"https://github.com/Lightricks/LTX-2"},{"provider":"LTX API","url":"https://console.ltx.io/playground/","docs":"https://docs.ltx.io"}],"capabilities":[{"name":"Open audio-video generation with multishot consistency","detail":"22B DiT that generates synchronized video and audio and keeps character, environment, lighting and voice consistent across cuts.","first":false,"discovered":"launch","source":"https://huggingface.co/Lightricks/LTX-2.5"},{"name":"Fast generation","detail":"10-second 720p image-to-video clip in 6.8 s on two GB200s (vendor claim); distilled variant uses 8 steps.","first":false,"discovered":"launch","source":"https://venturebeat.com/technology/ltx-2-5-can-generate-a-10-second-ai-video-from-an-image-in-just-6-8-seconds-on-nvidia-superchips-and-its-open-weights"},{"name":"Robotics / physical-AI checkpoint","detail":"Ships a pretrained checkpoint aimed at world-model use in robotics and physical AI.","first":false,"discovered":"launch","source":"https://venturebeat.com/technology/ltx-2-5-can-generate-a-10-second-ai-video-from-an-image-in-just-6-8-seconds-on-nvidia-superchips-and-its-open-weights"}],"entry":"2026-08-11-lightricks-ltx-2-5-open-weights","notes":"Full (trainable) and distilled DiT variants; Gemma 4 12B text encoder; up to 4K; frame counts must satisfy n % 8 == 1; ComfyUI native. Licence bans military/weapons use.","verified":"2026-10-03","body_md":"Lightricks' open-weights video+audio model; the most downloaded generative model on Hugging Face in early October 2026.","page_url":"https://postcutoff.com/m/ltx-2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":569,"you_url":null,"briefings":null},{"id":"nemotron-3-5-lightning","name":"NVIDIA Nemotron 3.5 Lightning (30B-A3B)","org":"NVIDIA","family":"Nemotron 3.5","released":"2026-08-11","status":"current","type":"llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"openmdw-1.1","context_window":1000000,"max_output":null,"knowledge_cutoff":"2025-09","pricing":{"input":0.06,"output":0.16,"unit":"per 1M tokens (USD) on OpenRouter (also :free variant)","source":"https://openrouter.ai/nvidia/nemotron-3.5-lightning"},"price_line":"$0.06 in, $0.16 out per 1M tokens","access":[{"provider":"NVIDIA API (build.nvidia.com)","model_id":"nvidia/nemotron-3.5-lightning-30b-a3b","endpoint":"https://integrate.api.nvidia.com/v1/chat/completions","docs":"https://build.nvidia.com"},{"provider":"OpenRouter","model_id":"nvidia/nemotron-3.5-lightning","url":"https://openrouter.ai/nvidia/nemotron-3.5-lightning"},{"provider":"Hugging Face","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16"}],"capabilities":[{"name":"Tiny-active MoE with 1M context","detail":"30B total / 3B active hybrid Mamba-2 + attention MoE with up to 1M context (256K on a single H100).","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16"},{"name":"Built for customization","detail":"Released with base checkpoint and NVFP4 builds (incl. speculative-decoding DSpark/DFlash variants); intended for fine-tuning and domain adaptation.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16"}],"entry":null,"notes":"Successor to Nemotron 3 Nano 30B-A3B (nvidia/nemotron-nano-3-30b-a3b on NIM). Knowledge cutoff = pre-training (Sep 2025); post-training to May 2026.","verified":"2026-09-29","body_md":"Fast, cheap small open model for high-throughput agents and fine-tuning.\n\n```bash\ncurl https://integrate.api.nvidia.com/v1/chat/completions -H \"Authorization: Bearer $NVIDIA_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"nvidia/nemotron-3.5-lightning-30b-a3b\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 , https://integrate.api.nvidia.com/v1/models","page_url":"https://postcutoff.com/m/nemotron-3-5-lightning/","events_after":834,"major_after":207,"historic_after":41,"missing_at_launch":261,"events_since_release":null,"you_url":"https://postcutoff.com/you/nemotron-3-5-lightning/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-09.md"}},{"id":"dyna-2","name":"DYNA-2 (World-Action Model)","org":"Dyna Robotics","family":"DYNA","released":"2026-08-10","status":"current","type":"robotics","modality_in":["image","video","text"],"modality_out":["video","action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Dyna Robotics (commercial deployments)","url":"https://www.dyna.co/dyna-2"}],"capabilities":[{"name":"Human-to-robot scaling law","detail":"Pretrained on 1M+ hours of egocentric human video (~170 years); on-robot normalized score rose from 20% to 53% across 14 tasks as pretraining scaled from 1k to 1M hours. Dyna calls it the first scaling law demonstrated across the embodiment gap.","first":true,"discovered":"launch","source":"https://www.dyna.co/dyna-2"},{"name":"World-action model","detail":"One video-diffusion (mixture-of-transformers, flow matching) model that denoises future video and an action chunk jointly or separately; one-step distilled video generation 90x faster than the teacher.","first":false,"discovered":"launch","source":"https://www.dyna.co/dyna-2"},{"name":"Production quality gains","detail":"87% zero-shot customer-quality pass rate at a customer deployment vs 46% for DYNA-1; 1.55x more task completions than DYNA-1; bottle-cap opening learned with 10 minutes of robot data.","first":false,"discovered":"launch","source":"https://www.prnewswire.com/news-releases/dyna-robotics-unveils-dyna-2-world-action-model-demonstrating-first-true-scaling-law-in-robotics-powered-entirely-by-human-data-302847114.html"}],"entry":"2026-08-10-dyna-robotics-dyna-2","notes":"Predecessor DYNA-1 (2025) runs in production in hotels, restaurants and laundromats (towel folding etc.). No API or weights; adapts to arms, humanoid prototypes and dexterous hands with hours of local fine-tuning. Figure (Helix 2.5), Generalist (GEN-1) and Dyna all reported human-video scaling in 2026, so 'first' claims overlap. Company-reported.","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/dyna-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":573,"you_url":null,"briefings":null},{"id":"soniox-tts-v2","name":"Soniox TTS v2","org":"Soniox","family":"Soniox TTS","released":"2026-08-10","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":4,"output":21.5,"unit":"USD per 1M tokens (text in / audio out); ≈ $0.70 per hour of generated speech (1 hour ≈ 30,000 audio tokens)","source":"https://soniox.com/pricing"},"price_line":"$4 in, $21.50 out per 1M tokens","access":[{"provider":"Soniox API (real-time streaming, WebSocket)","model_id":"tts-rt-v2","docs":"https://soniox.com/text-to-speech"}],"capabilities":[{"name":"60+ languages in one model, mid-sentence switching","detail":"Single multilingual model with mixed-language text and mid-sentence language switching; Soniox claims 'hallucination-free' output (no invented or dropped words) and accurate reading of emails, phone numbers and IDs.","first":false,"discovered":"launch","source":"https://soniox.com/blog/soniox-text-to-speech"},{"name":"Audio tags and 20-second voice cloning (v2)","detail":"TTS v2 adds expressive audio tags (whispering, laughter, hesitation, excitement), voice cloning from ~20 s of reference audio, and character-level timestamps.","first":false,"discovered":"launch","source":"https://audioxpress.com/news/soniox-tts-v2-adds-expressive-control-and-voice-cloning-to-its-multilingual-voice-ai-platform"}],"entry":null,"notes":"Soniox launched TTS on 2026-04-23 (tts-rt-v1); TTS v2 (tts-rt-v2, replacing v1) was reported by audioXpress on 2026-08-10. Streaming only; regions US, EU, Japan. The v2 date is from secondary press, not a Soniox post.","verified":"2026-09-29","body_md":"Soniox's text-to-speech, the companion to its [v5 STT](soniox-stt-v5.md), aimed at multilingual voice agents.\n\nSources: https://soniox.com/blog/soniox-text-to-speech , https://soniox.com/pricing , https://audioxpress.com/news/soniox-tts-v2-adds-expressive-control-and-voice-cloning-to-its-multilingual-voice-ai-platform","page_url":"https://postcutoff.com/m/soniox-tts-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":573,"you_url":null,"briefings":null},{"id":"grok-imagine-image-2","name":"Grok Imagine Image 2.0","org":"xAI","family":"Grok Imagine","released":"2026-08-07","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.04,"unit":"per output image (base tier per xAI models page)","source":"https://docs.x.ai/docs/models"},"price_line":"$0.04 per image","access":[{"provider":"xAI API","model_id":"grok-imagine-image-2.0","endpoint":"https://api.x.ai/v1/images/generations","docs":"https://docs.x.ai/docs/guides/image-generation"},{"provider":"xAI API (edits)","model_id":"grok-imagine-image-2.0","endpoint":"https://api.x.ai/v1/images/edits"},{"provider":"Web app","url":"https://grok.com"}],"capabilities":[{"name":"Generation + editing in one model","detail":"Text-to-image and image editing (URL or base64 input) via /v1/images/generations and /v1/images/edits.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/guides/image-generation"},{"name":"Top-2 on Arena image leaderboards at launch","detail":"xAI reported #2 on both Arena Text-to-Image and Arena Image Edit at launch (Aug 7, 2026).","first":false,"discovered":"launch","source":"https://kie.ai/blog/grok-imagine-image-2-0-release"}],"entry":null,"notes":"xAI's recommended image model; cheaper grok-imagine-image ($0.02) and grok-imagine-image-quality ($0.05) also listed. App launch 2026-08-07, API shortly after. Third-party reports of resolution/quality price tiers not verified on official page.","verified":"2026-09-29","body_md":"Image generation and editing for Grok / Imagine apps and API.\n\n```bash\ncurl -X POST https://api.x.ai/v1/images/generations -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-imagine-image-2.0\",\"prompt\":\"A collage of London landmarks in stenciled street-art style\"}'\n```\n\nSources: https://docs.x.ai/docs/models , https://docs.x.ai/docs/guides/image-generation","page_url":"https://postcutoff.com/m/grok-imagine-image-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":576,"you_url":null,"briefings":null},{"id":"qwen3-8-27b","name":"Qwen3.8-27B","org":"Alibaba (Qwen)","family":"Qwen3.8","released":"2026-08-05","status":"current","type":"multimodal","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/Qwen/Qwen3.8-27B"},{"provider":"Hugging Face (FP8)","url":"https://huggingface.co/Qwen/Qwen3.8-27B-FP8"},{"provider":"OpenRouter","model_id":"qwen/qwen3.8-27b","url":"https://openrouter.ai/qwen/qwen3.8-27b"},{"provider":"OpenRouter (free tier)","model_id":"qwen/qwen3.8-27b:free","url":"https://openrouter.ai/qwen/qwen3.8-27b"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"Dense open VLM with agentic focus","detail":"27B dense native vision-language model (images and hour-scale video) tuned for coding and long-horizon agent tasks, Apache-2.0.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-27B"},{"name":"Thinking control","detail":"Thinking on by default, can be disabled per request; reasoning_effort and preserve_thinking supported.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-27B"},{"name":"Extensible to 1M context","detail":"262,144 tokens native, extensible up to 1,000,000.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-27B"}],"entry":null,"notes":"Best Apache-2.0 Qwen for self-hosting; also the go-to open Qwen VL model (Qwen3-VL successor). First-party hosted API 'coming soon' on Qwen Cloud at time of check. Pricing not verified (no first-party price). ARC Prize verified (Oct 1 2026): ARC-AGI-2 42.4% at $0.45/task, ARC-AGI-1 87.5% at $0.22/task (x.com/arcprize/status/2105751611570069908).","verified":"2026-09-29","body_md":"Open-weight (Apache-2.0) dense multimodal model for local/self-hosted coding agents and vision tasks; runs on vLLM, SGLang, Transformers.\n\n```bash\nvllm serve Qwen/Qwen3.8-27B --max-model-len 262144\n```\n\nSources: https://huggingface.co/Qwen/Qwen3.8-27B · https://openrouter.ai/qwen/qwen3.8-27b","page_url":"https://postcutoff.com/m/qwen3-8-27b/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":579,"you_url":"https://postcutoff.com/you/qwen3-8-27b/","briefings":null},{"id":"seedrealtime","name":"SeedRealtime (Doubao realtime audio-visual model)","org":"ByteDance","family":"Seed","released":"2026-08-05","status":"current","type":"audio/speech","modality_in":["audio","video","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Web app (Doubao / Dola)","url":"https://dola.com/chat"},{"provider":"BytePlus Playground","url":"https://ai.byteplus.com/en/playground"}],"capabilities":[{"name":"Native audio-visual full-duplex LLM","detail":"Single end-to-end model perceives continuous audio, video and text streams while listening and speaking (no ASR/VLM/TTS cascade); resolves homophones from visual context and temporal references to what it sees.","first":false,"discovered":"launch","source":"https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction"},{"name":"Proactive turn-taking","detail":"ByteDance says it halves audio-visual conversational pacing problems vs cascaded systems (fewer cut-offs, slow replies, false triggers) and can speak up proactively.","first":false,"discovered":"launch","source":"https://seed.bytedance.com/en/SeedRealtime"}],"entry":"2026-08-05-bytedance-seedrealtime","notes":"Deployed at scale in the Doubao app (Dola internationally). No public API model id, pricing or benchmark numbers published; Volcengine offers a separate Doubao end-to-end realtime dialogue API (/api/v3/realtime/dialogue) whose relation to SeedRealtime is unverified. Some press calls it the first model to watch, listen and speak simultaneously; not claimed by ByteDance, and Gemini Live / GPT-Realtime already accepted video.","verified":null,"body_md":"Sources: https://seed.bytedance.com/en/SeedRealtime · https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction · https://technode.com/2026/08/05/bytedance-launches-seedrealtime-full-duplex-audio-video-model/","page_url":"https://postcutoff.com/m/seedrealtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":579,"you_url":null,"briefings":null},{"id":"nemotron-voicechat","name":"NVIDIA NemotronLabs VoiceChat 11B (and PersonaPlex-7B)","org":"NVIDIA","family":"Nemotron Speech","released":"2026-08-03","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["audio","text"],"open_weights":true,"model_license":"OpenMDW-1.1","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/NVIDIA-NemotronLabs-VoiceChat-11B","url":"https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B"},{"provider":"Hugging Face","model_id":"nvidia/personaplex-7b-v1","url":"https://huggingface.co/nvidia/personaplex-7b-v1"},{"provider":"arXiv","url":"https://arxiv.org/abs/2609.21967"}],"capabilities":[{"name":"Open full-duplex speech model with tool calling","detail":"End-to-end (FastConformer encoder + Nemotron Nano v2 9B + TTS decoder, 11B total) full-duplex voice chat that calls tools mid-conversation; NVIDIA calls it the first open full-duplex model to support tool calling. BFCL-v3 (AU Harness) 56.1%, Full-Duplex-Bench v3 tool selection 82.5%.","first":true,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B"},{"name":"Natural turn-taking","detail":"~450 ms turn-taking latency; #2 among open models on VoiceBench and Full-Duplex-Bench 1.0 (smooth turn-taking 0.82, interruption latency 480 ms).","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B"},{"name":"PersonaPlex: persona + voice prompted full duplex","detail":"PersonaPlex-7B-v1 (2026-01-15), fine-tuned from Kyutai Moshiko, takes a voice prompt and a text persona/role prompt.","first":false,"discovered":"later","source":"https://huggingface.co/nvidia/personaplex-7b-v1"}],"entry":"2026-08-03-nvidia-nemotronlabs-voicechat","notes":"English only. Requires datacenter GPU (A100/H100/H200/B100/B200 or RTX 6000). 'First' is NVIDIA's claim on the model card. HF card release date 2026-08-03; arXiv paper 2609.21967 (Sept 2026).","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B , https://huggingface.co/nvidia/personaplex-7b-v1","page_url":"https://postcutoff.com/m/nemotron-voicechat/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":588,"you_url":null,"briefings":null},{"id":"glm-5-3","name":"GLM-5.3","org":"Zhipu AI (Z.ai)","family":"GLM-5","released":"2026-08","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"glm-5.3 (custom)","context_window":1000000,"max_output":128000,"knowledge_cutoff":null,"pricing":{"input":1.4,"output":4.4,"cache_read":0.26,"unit":"per 1M tokens (USD)","source":"https://docs.z.ai/guides/overview/pricing"},"price_line":"$1.40 in, $4.40 out per 1M tokens","access":[{"provider":"Z.ai API","model_id":"glm-5.3","endpoint":"https://api.z.ai/api/paas/v4","docs":"https://docs.z.ai/guides/llm/glm-5.3"},{"provider":"Z.ai API (Anthropic format)","model_id":"glm-5.3","endpoint":"https://api.z.ai/api/anthropic","docs":"https://docs.z.ai/guides/llm/glm-5.3"},{"provider":"Alibaba Cloud Model Studio","model_id":"ZHIPU/GLM-5.3","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"OpenRouter","model_id":"z-ai/glm-5.3","url":"https://openrouter.ai/z-ai/glm-5.3"},{"provider":"Hugging Face","url":"https://huggingface.co/zai-org/GLM-5.3"},{"provider":"Web app","url":"https://chat.z.ai"}],"capabilities":[{"name":"Post-training-only jump in coding","detail":"Same base as GLM-5.2; Z.ai reports +50% on its Code Bench and open-model SOTA on Terminal Bench 3.0 and Agents' Last Exam (CLI).","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/llm/glm-5.3"},{"name":"Emergent cyber capability","detail":"Best CyberGym vulnerability-discovery score to date per Z.ai; exploitation benchmark scores more than double GLM-5.2's.","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/llm/glm-5.3"},{"name":"Always-on reasoning with effort levels","detail":"thinking.type disabled no longer allowed; reasoning_effort low/high/max (default max).","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/llm/glm-5.3"},{"name":"Coding Plan integration","detail":"Available in the GLM Coding Plan (points-based; off-peak/weekend calls cost 50% points) for Claude Code, Cline, OpenCode etc.","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/llm/glm-5.3"}],"entry":null,"notes":"Z.ai flagship. Text-only input. Migration: requests with thinking disabled fail - set enabled + reasoning_effort low. Coding Plan base URL is https://api.z.ai/api/coding/paas/v4. Release day not verified (OpenRouter 2026-08-18, HF 2026-08-25). GLM-5.2 (same price, MIT weights) still listed.","verified":"2026-09-29","body_md":"Z.ai's top model for long-horizon software engineering, terminal agents and security research; 1M context, 128K output.\n\n```bash\ncurl https://api.z.ai/api/paas/v4/chat/completions \\\n -H \"Authorization: Bearer $ZAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"glm-5.3\",\"thinking\":{\"type\":\"enabled\"},\"reasoning_effort\":\"max\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.z.ai/guides/llm/glm-5.3 · https://docs.z.ai/guides/overview/pricing · https://huggingface.co/zai-org/GLM-5.3","page_url":"https://postcutoff.com/m/glm-5-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":508,"you_url":"https://postcutoff.com/you/glm-5-3/","briefings":null},{"id":"glm-5-3-flash","name":"GLM-5.3-Flash / FlashX","org":"Zhipu AI (Z.ai)","family":"GLM-5","released":"2026-08","status":"current","type":"multimodal","modality_in":["text","image","video","pdf"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":1000000,"max_output":128000,"knowledge_cutoff":null,"pricing":{"input":0.15,"output":0.5,"cache_read":0.03,"unit":"per 1M tokens (USD) for glm-5.3-flash; glm-5.3-flashx (~200 tok/s): 0.37 in / 1.25 out / 0.075 cached","source":"https://docs.z.ai/guides/overview/pricing"},"price_line":"$0.15 in, $0.50 out per 1M tokens","access":[{"provider":"Z.ai API","model_id":"glm-5.3-flash","endpoint":"https://api.z.ai/api/paas/v4","docs":"https://docs.z.ai/guides/vlm/glm-5.3-flash"},{"provider":"Z.ai API (fast)","model_id":"glm-5.3-flashx","endpoint":"https://api.z.ai/api/paas/v4","docs":"https://docs.z.ai/guides/vlm/glm-5.3-flash"},{"provider":"OpenRouter","model_id":"z-ai/glm-5.3-flash","url":"https://openrouter.ai/z-ai/glm-5.3-flash"},{"provider":"OpenRouter (FlashX)","model_id":"z-ai/glm-5.3-flashx","url":"https://openrouter.ai/z-ai/glm-5.3-flashx"},{"provider":"RunInfra (third-party inference, also via Vercel AI Gateway)","url":"https://runinfra.ai/inference-api/glm-5-3-flash","docs":"https://x.com/runinfrai/status/2105186912634057086","notes":"Vendor claims (2026-09-30): rewritten FP8 kernels, ~670 tok/s on Vercel AI Gateway, $0.11 in / $0.45 out / $0.03 cached per 1M tokens, 1M context, runs on NVIDIA and AMD, OpenAI- and Anthropic-compatible APIs, zero data retention. Not independently benchmarked."},{"provider":"Hugging Face","url":"https://huggingface.co/zai-org/GLM-5.3-Flash"},{"provider":"Web app","url":"https://chat.z.ai"}],"capabilities":[{"name":"First native multimodal GLM-5 model","detail":"First GLM-5-series model with native vision (image, video, file input); vision used inside the coding loop (UI replication, Blender, browser/computer-use agents).","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/vlm/glm-5.3-flash"},{"name":"Sparse + linear attention hybrid","detail":"320B total / 18B active; Z.ai claims it is the first open-source frontier model combining sparse and linear attention (3.01x less attention compute, 4.44x smaller KV cache vs GLM-5.3).","first":true,"discovered":"launch","source":"https://docs.z.ai/guides/vlm/glm-5.3-flash"},{"name":"Office deliverables with visual self-check","detail":"Produces PPTX/PDF/DOCX/XLSX and renders them to catch overflow and layout issues.","first":false,"discovered":"launch","source":"https://docs.z.ai/guides/vlm/glm-5.3-flash"}],"entry":null,"notes":"Z.ai says it beats GLM-5.2 at a fraction of the cost; 3x Coding Plan quota vs GLM-5.3 (FlashX not yet on the plan). Thinking cannot be disabled. 'first' claim is the vendor's own. ARC Prize verified (Sept 29 2026): ARC-AGI-2 65.8% at $0.09/task, ARC-AGI-1 91.0% at $0.04/task; it was tested earlier in stealth on OpenRouter as Ox Alpha (x.com/arcprize/status/2104996180425884119).","verified":"2026-09-30","body_md":"Cheap multimodal GLM for visual coding, agents and office documents; open weights under MIT.\n\n```bash\ncurl https://api.z.ai/api/paas/v4/chat/completions \\\n -H \"Authorization: Bearer $ZAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"glm-5.3-flash\",\"messages\":[{\"role\":\"user\",\"content\":[{\"type\":\"image_url\",\"image_url\":{\"url\":\"https://example.com/ui.png\"}},{\"type\":\"text\",\"text\":\"Rebuild this UI in React\"}]}]}'\n```\n\nSources: https://docs.z.ai/guides/vlm/glm-5.3-flash · https://docs.z.ai/guides/overview/pricing · https://huggingface.co/zai-org/GLM-5.3-Flash","page_url":"https://postcutoff.com/m/glm-5-3-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":508,"you_url":"https://postcutoff.com/you/glm-5-3-flash/","briefings":null},{"id":"muse-glimmer-30b","name":"Muse Glimmer 30B","org":"Meta","family":"Muse","released":"2026-08","status":"current","type":"llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":131072,"max_output":null,"knowledge_cutoff":"2026-01","pricing":{"input":0.3,"output":1.2,"unit":"per 1M tokens (USD) on OpenRouter; open weights free to self-host","source":"https://openrouter.ai/meta/muse-glimmer-30b"},"price_line":"$0.30 in, $1.20 out per 1M tokens","access":[{"provider":"Hugging Face","url":"https://huggingface.co/meta-models/Muse-Glimmer-30B"},{"provider":"OpenRouter","model_id":"meta/muse-glimmer-30b","url":"https://openrouter.ai/meta/muse-glimmer-30b"}],"capabilities":[{"name":"Meta open weights under Apache 2.0","detail":"~29.6B dense text+image model released Apache 2.0 (Llama used a custom community license), with llama.cpp / MLX / ExecuTorch integrations.","first":false,"discovered":"launch","source":"https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"},{"name":"Local agents on one consumer GPU","detail":"Quantized to under 20GB for 24-32GB consumer GPUs/Macs; bundled DFlash drafter for speculative decoding gives ~3.1x speed-up on RTX 5090.","first":false,"discovered":"launch","source":"https://huggingface.co/meta-models/Muse-Glimmer-30B"},{"name":"Agentic focus for its size","detail":"Optimized for multi-step reasoning, reliable tool use and failure recovery; Meta benchmarks it as competitive with Gemma4-31B and Qwen3.6-27B on agentic/coding evals.","first":false,"discovered":"launch","source":"https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model"}],"entry":null,"notes":"Released early Aug 2026 (exact day not verified). HF org is meta-models, not meta-llama. No first-party Meta API id verified.","verified":"2026-09-29","body_md":"Open-weight agentic/vision model for local deployment.\n\n```python\nfrom transformers import pipeline\npipe = pipeline(\"image-text-to-text\", model=\"meta-models/Muse-Glimmer-30B\")\n```\n\nSources: https://huggingface.co/meta-models/Muse-Glimmer-30B , https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model","page_url":"https://postcutoff.com/m/muse-glimmer-30b/","events_after":793,"major_after":197,"historic_after":40,"missing_at_launch":198,"events_since_release":null,"you_url":"https://postcutoff.com/you/muse-glimmer-30b/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-01.md"}},{"id":"qwen3-8-flash","name":"Qwen3.8-Flash","org":"Alibaba (Qwen)","family":"Qwen3.8","released":"2026-08","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"qwen-community-1.0 (open weights Qwen3.8-Flash-Next)","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.15,"output":0.47,"unit":"per 1M tokens (USD), Singapore/International region, input up to 1M","source":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},"price_line":"$0.15 in, $0.47 out per 1M tokens","access":[{"provider":"Alibaba Cloud Model Studio (DashScope, Singapore/Intl)","model_id":"qwen3.8-flash","endpoint":"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash"},{"provider":"OpenRouter","model_id":"qwen/qwen3.8-flash","url":"https://openrouter.ai/qwen/qwen3.8-flash"},{"provider":"Hugging Face (Qwen3.8-Flash-Next, base of the API model)","url":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next"},{"provider":"Local, community (Strata one-click installer for Qwen3.8-Flash-Next GGUF quants; 12 GB+ GPU, 32 GB+ RAM)","url":"https://github.com/Niko1221/Strata","endpoint":"http://127.0.0.1:8080/v1"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"Preview of the Qwen4 architecture","detail":"Built on Qwen3.8-Flash-Next, an experimental preview of the architecture that will underpin Qwen4 (Gated DeltaNet + Qwen Sparse Attention, Gated Residual, N-gram Embedding).","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next"},{"name":"Block-level sparse attention (QSA)","detail":"Qwen Sparse Attention selects micro-blocks rather than tokens, cutting long-context latency for agentic workloads.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next"},{"name":"OpenAI + Anthropic protocol compatibility","detail":"Works directly with Claude Code and Codex; 1M context, image/video understanding, desktop-app operation.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash"}],"entry":null,"notes":"Low-cost default in Model Studio (maps to 'GPT-5.4-mini / Haiku 4.5' tier per Alibaba). Max output not verified. Release day not verified (OpenRouter listing 2026-08-26).","verified":"2026-09-29","body_md":"Cheap, fast multimodal workhorse for coding assistants, agents and high-concurrency apps; 1M context.\n\n```bash\ncurl \"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions\" \\\n -H \"Authorization: Bearer $DASHSCOPE_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"qwen3.8-flash\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-flash · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://huggingface.co/Qwen/Qwen3.8-Flash-Next","page_url":"https://postcutoff.com/m/qwen3-8-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":508,"you_url":"https://postcutoff.com/you/qwen3-8-flash/","briefings":null},{"id":"qwen3-8-max","name":"Qwen3.8-Max","org":"Alibaba (Qwen)","family":"Qwen3.8","released":"2026-08","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"qwen3.8-max (custom, for open weights Qwen3.8-2.4T-A95B)","context_window":1000000,"max_output":131072,"knowledge_cutoff":null,"pricing":{"input":2,"output":6,"cache_read":0.25,"unit":"per 1M tokens (USD), Singapore/International region, input up to 1M; Beijing/Global regions 1.65/4.951","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max"},"price_line":"$2 in, $6 out per 1M tokens","access":[{"provider":"Alibaba Cloud Model Studio (DashScope, Singapore/Intl)","model_id":"qwen3.8-max","endpoint":"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max"},{"provider":"Alibaba Cloud Model Studio (US Virginia)","model_id":"qwen3.8-max","endpoint":"https://dashscope-us.aliyuncs.com/compatible-mode/v1","docs":"https://www.alibabacloud.com/help/en/model-studio/compatibility-of-openai-with-dashscope"},{"provider":"OpenRouter","model_id":"qwen/qwen3.8-max-0902","url":"https://openrouter.ai/qwen/qwen3.8-max-0902"},{"provider":"OpenRouter (open-weight base)","model_id":"qwen/qwen3.8-2.4t-a95b","url":"https://openrouter.ai/qwen/qwen3.8-2.4t-a95b"},{"provider":"Hugging Face","url":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"First open-weight Qwen-Max-class model","detail":"Qwen3.8 brings a Max-class model to open release for the first time (Qwen3.8-2.4T-A95B, 2.4T total / 95B active MoE); the API version adds vision input, non-thinking mode, 1M context and built-in tools.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"},{"name":"Multi-day autonomous coding","detail":"Alibaba markets it as able to code autonomously for over ten days to deliver complete projects, with closed-loop planning and iteration.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max"},{"name":"Native vision in the agent loop","detail":"Image and video understanding used throughout planning, execution and verification; parses ultra-long documents and long videos.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max"},{"name":"Tunable and preserved thinking","detail":"reasoning_effort controls depth; preserve_thinking keeps reasoning context from earlier turns.","first":false,"discovered":"launch","source":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"}],"entry":null,"notes":"Alibaba's top model. Apsara 2026 (2026-09-22): Alibaba says an updated Qwen3.8-Max went through 33 automated self-improvement cycles, raising its Artificial Analysis score from 40 to 45 (company claim, https://www.alibabacloud.com/en/press-room/alibaba-unveils-roadmap-on-full-stack-ai-strategy). Snapshot qwen3.8-max-0902; fast tier qwen3.8-max-prime (OpenRouter qwen/qwen3.8-max-prime, Beijing 3.301/9.902). Singapore endpoint needs your WorkspaceId (old dashscope-intl domain is being migrated). Also sold via Qwen Cloud (qwencloud.com). Release day not verified (weights on HF 2026-08-08). Knowledge cutoff not published.","verified":"2026-09-29","body_md":"Alibaba's flagship for hard reasoning, long-horizon coding and professional work (law, finance, design); 1M context, text/image/video in.\n\n```bash\ncurl \"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions\" \\\n -H \"Authorization: Bearer $DASHSCOPE_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"qwen3.8-max\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}],\"enable_thinking\":true}'\n```\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen3-8-max · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B","page_url":"https://postcutoff.com/m/qwen3-8-max/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":508,"you_url":"https://postcutoff.com/you/qwen3-8-max/","briefings":null},{"id":"minimax-h3","name":"MiniMax H3","org":"MiniMax","family":"MiniMax H (Hailuo successor)","released":"2026-07-31","status":"current","type":"video-gen","modality_in":["text","image","video","audio"],"modality_out":["video","audio"],"open_weights":true,"model_license":"minimax-h3-community-license-agreement","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second_768p":0.08,"per_second_2k":0.13,"unit":"per second of output video (USD). H3-Max (fal.ai post-trained, fast): 480P 0.05/s, 768P 0.08/s. Extra input images 0.04 each after 5 free","source":"https://platform.minimax.io/docs/guides/pricing-paygo"},"price_line":"$0.08 per second, 768p","access":[{"provider":"MiniMax API (Video Generation V2)","model_id":"MiniMax-H3","endpoint":"https://api.minimax.io","docs":"https://platform.minimax.io/docs/api-reference/video-generation-v2-create"},{"provider":"MiniMax API (fast variant)","model_id":"MiniMax-H3-Max","endpoint":"https://api.minimax.io","docs":"https://platform.minimax.io/docs/api-reference/video-generation-v2-create"},{"provider":"Hugging Face","url":"https://huggingface.co/MiniMaxAI/MiniMax-H3"}],"capabilities":[{"name":"Open omni-modal video model with native audio","detail":"Understands mixed text/image/video/audio context and generates video with native stereo audio, up to 2K and 15 s.","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-H3"},{"name":"H3-Context-IR prompt pipeline","detail":"Hosted system turns free-form multimodal instructions into a structured intermediate representation before generation (API-only, not open-sourced).","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-H3"},{"name":"768P to 2K regeneration","detail":"H3-Regenerate-2K re-renders a 768P result with the original context into 2K (0.05 USD/s).","first":false,"discovered":"launch","source":"https://platform.minimax.io/docs/guides/pricing-paygo"}],"entry":null,"notes":"Replaces Hailuo 2.3 / 2.3-Fast / 02 (now legacy: e.g. MiniMax-Hailuo-2.3 0.28 USD per 768P 6s clip). Modes: T2V, I2V, first/last frame, multimodal reference; 4-15 s, 24 fps. Open release is full-attention only.","verified":"2026-09-29","body_md":"MiniMax's current video generator (successor to Hailuo): text/image/video/audio-conditioned clips with sound, up to 2K. Async task API (create task, then query by task_id).\n\nSources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://huggingface.co/MiniMaxAI/MiniMax-H3 · https://platform.minimax.io/docs/release-notes/models","page_url":"https://postcutoff.com/m/minimax-h3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":595,"you_url":null,"briefings":null},{"id":"gemini-robotics-2","name":"Gemini Robotics 2","org":"Google DeepMind","family":"Gemini Robotics","released":"2026-07-30","status":"preview","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Gemini Robotics trusted tester / early-access program (application form)","url":"https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform","docs":"https://deepmind.google/models/gemini-robotics/"}],"capabilities":[{"name":"Whole-body humanoid control from a VLA","detail":"Google's first VLA to control an entire humanoid (walking, crouching, balancing while manipulating) rather than only the upper body; e.g. Apollo with Inspire hands: 68.4% pick from table, 45.7% from floor, 76.3% from shelf.","first":false,"discovered":"launch","source":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"},{"name":"Multi-finger and gripper dexterity across embodiments","detail":"Same model drives multi-fingered hands and grippers (Franka Duo: 89.6% precise insertion; Apollo with SharpaWave hands: 92% unscrew bulb).","first":false,"discovered":"launch","source":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"},{"name":"Paired with ER 2 planner","detail":"Designed to be called by Gemini Robotics ER 2, which plans, tracks progress and coordinates multiple robots.","first":false,"discovered":"launch","source":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"}],"entry":"2026-07-30-gemini-robotics-2","notes":"Vision-language-action model (outputs robot motor commands; modality 'action'). No public API or weights: available only to early-access partners (Apptronik, Boston Dynamics, Agile Robots, Franka, 100+ trusted testers) via waitlist form. DeepMind says 'for the first time, our model can control entire humanoid robots' - first for Google, not industry-first (Figure Helix 02 showed whole-body VLA control in Jan 2026). Predecessor: Gemini Robotics 1.5 (Sep 2025), itself trusted-tester only.","verified":"2026-09-29","body_md":"Google DeepMind's flagship vision-language-action model for humanoids and bi-arm robots.\n\nSources: [DeepMind blog](https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/), [Gemini Robotics page](https://deepmind.google/models/gemini-robotics/).","page_url":"https://postcutoff.com/m/gemini-robotics-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":596,"you_url":null,"briefings":null},{"id":"gemini-robotics-er-2","name":"Gemini Robotics ER 2","org":"Google DeepMind","family":"Gemini Robotics","released":"2026-07-30","status":"preview","type":"robotics","modality_in":["text","image","video","audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":1,"output":5,"unit":"per 1M tokens (text/image/video/audio input); introductory rate through 2026-12-31, rising to $2.00 in / $10.00 out from 2027-01-01; Batch API half price","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1 in, $5 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-robotics-er-2-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-robotics-er-2-preview:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview"},{"provider":"Gemini API (Live API, streaming)","model_id":"gemini-robotics-er-2-streaming-preview","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-streaming-preview"},{"provider":"Google AI Studio","url":"https://aistudio.google.com"},{"provider":"Gemini Enterprise Agent Platform (Google Cloud, private preview)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/gemini-robotics-er"},{"provider":"Sample code (GitHub)","url":"https://github.com/google-gemini/robotics-samples"}],"capabilities":[{"name":"Embodied reasoning \"robot brain\" in a public API","detail":"Spatial reasoning (points, boxes, trajectories), multi-step task planning, tool/function calling and code execution to orchestrate a robot's VLA or controller; publicly callable, unlike the VLA models.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/robotics-overview"},{"name":"Continuous video monitoring and task-progress tracking","detail":"Watches video feeds to track progress and adapt; Google reports 91.3% moment-finding accuracy (0.96 s mean absolute distance) at ~4x the speed of the previous generation and 57.4% progress classification.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/"},{"name":"Low-latency streaming via Live API","detail":"Separate gemini-robotics-er-2-streaming-preview id supports bidirectional audio/video streaming with function calling and thinking (no caching, code execution or structured output).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/robotics-streaming"},{"name":"Multi-robot collaboration","detail":"Coordinates heterogeneous robots (e.g. wheeled rovers and humanoids, Boston Dynamics Spot demo) to communicate and hand off tasks.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/"}],"entry":"2026-07-30-gemini-robotics-2","notes":"Vision-language model for robotics (outputs text/JSON, not motor commands). 131,072 input / 65,536 output tokens. Standard id supports caching, code execution, computer use, file search, function calling, Search and Maps grounding, structured outputs and thinking. Replaces gemini-robotics-er-1.6-preview (shut down 2026-08-31). No GA id yet. Knowledge cutoff not stated.","verified":"2026-09-29","body_md":"The hosted, publicly callable half of the Gemini Robotics 2 stack: a high-level planner that points at objects, plans multi-step tasks and calls a robot's own VLA/skills as tools.\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [Google blog](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/), [model card](https://deepmind.google/models/model-cards/gemini-robotics-er-2/).","page_url":"https://postcutoff.com/m/gemini-robotics-er-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":596,"you_url":null,"briefings":null},{"id":"gemini-robotics-on-device-2","name":"Gemini Robotics On-Device 2","org":"Google DeepMind","family":"Gemini Robotics","released":"2026-07-30","status":"preview","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Gemini Robotics trusted tester / early-access program (application form)","url":"https://docs.google.com/forms/d/1sM5GqcVMWv-KmKY3TOMpVtQ-lDFeAftQ-d9xQn92jCE/viewform","docs":"https://deepmind.google/models/gemini-robotics/"}],"capabilities":[{"name":"Local VLA inference on robot hardware","detail":"Lightweight version of the Gemini Robotics VLA optimized to run locally without a network connection.","first":false,"discovered":"launch","source":"https://deepmind.google/models/gemini-robotics/"},{"name":"Fast adaptation to new embodiments","detail":"Adapts to completely new robot bodies with a few hours of data; typically fewer than 200 examples for a new bi-arm robot (uses motion transfer from Gemini Robotics 1.5).","first":false,"discovered":"launch","source":"https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/"}],"entry":"2026-07-30-gemini-robotics-2","notes":"Successor to Gemini Robotics On-Device (June 2025). Trusted-tester / partner access only; parameter count and hardware requirements not published.","verified":"2026-09-29","body_md":"On-robot VLA for low-latency or offline deployments.\n\nSources: [DeepMind blog](https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/).","page_url":"https://postcutoff.com/m/gemini-robotics-on-device-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":596,"you_url":null,"briefings":null},{"id":"grok-voice-think-fast-2-0","name":"Grok Voice Think Fast 2.0","org":"xAI","family":"Grok Voice","released":"2026-07-29","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.08,"unit":"USD per minute of audio ($4.80/hr), plus $0.004 per text input (as listed on the xAI pricing page)","source":"https://docs.x.ai/developers/pricing"},"price_line":"$0.08 per minute","access":[{"provider":"xAI API (Voice Agent / speech-to-speech, WebSocket)","model_id":"grok-voice-think-fast-2.0","endpoint":"wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0","docs":"https://docs.x.ai/developers/model-capabilities/audio/voice-agent"},{"provider":"xAI API (alias)","model_id":"grok-voice-latest","docs":"https://docs.x.ai/developers/model-capabilities/audio/voice-agent"},{"provider":"Web app","url":"https://grok.com"}],"capabilities":[{"name":"Reasoning while speaking","detail":"Speech-to-speech model that reasons in real time (reasoning effort 'high' by default, can be set to 'none'); 97.2% Big Bench Audio, 82.9 on the Artificial Analysis Speech-to-Speech Quality Index (vs 75.7 for v1.0).","first":false,"discovered":"launch","source":"https://x.ai/news/grok-voice-think-fast-2"},{"name":"Faster first audio","detail":"Time to first audio cut from 1.25 s (v1.0) to 0.70 s; Full Duplex Bench 95.1%, tau-voice Bench 56.5% (xAI-reported).","first":false,"discovered":"launch","source":"https://x.ai/news/grok-voice-think-fast-2"},{"name":"Built-in server-side tools","detail":"Web search, X search, collections (file) search and remote MCP callable from inside a voice session, plus custom functions.","first":false,"discovered":"launch","source":"https://docs.x.ai/developers/model-capabilities/audio/voice-agent"},{"name":"OpenAI Realtime-compatible protocol","detail":"Largely compatible with the OpenAI Realtime SDK: change base URL to https://api.x.ai/v1 and the API key (minor event-name differences).","first":false,"discovered":"launch","source":"https://docs.x.ai/developers/model-capabilities/audio/voice-agent"}],"entry":"2026-07-29-grok-voice-think-fast-2","notes":"Released 2026-07-29; grok-voice-latest switched to it on 2026-08-05. Predecessor grok-voice-think-fast-1.0 can still be pinned. 20+ languages; audio PCM (8-48 kHz), Opus 24 kHz, G.711 mu-law/A-law; server VAD, session resumption (30 min), custom cloned voices. xAI says Starlink A/B tests raised sales conversion and support containment. Benchmarks are xAI-reported.","verified":"2026-09-29","body_md":"xAI's realtime voice-agent model (also used in the Grok app, Tesla vehicles and Starlink support).\n\n```python\n# OpenAI Realtime-compatible WebSocket\nimport websockets, os\nurl = \"wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0\"\nws = await websockets.connect(url, additional_headers={\"Authorization\": f\"Bearer {os.environ['XAI_API_KEY']}\"})\n```\n\nSources: https://x.ai/news/grok-voice-think-fast-2 · https://docs.x.ai/developers/model-capabilities/audio/voice-agent · https://docs.x.ai/developers/pricing","page_url":"https://postcutoff.com/m/grok-voice-think-fast-2-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":600,"you_url":null,"briefings":null},{"id":"lyria-3-5","name":"Lyria 3.5","org":"Google DeepMind","family":"Lyria","released":"2026-07-29","status":"current","type":"music","modality_in":["text","image"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_song":0.08,"unit":"per full song (~2 min). Lyria 3 Clip (30 s, lyria-3-clip-preview) $0.04 per clip; lyria-3-pro-preview $0.08 per song","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.08 per song","access":[{"provider":"Gemini API (Interactions API)","model_id":"lyria-3.5","endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","docs":"https://ai.google.dev/gemini-api/docs/music-generation"},{"provider":"Gemini API (Interactions API)","model_id":"lyria-3-clip-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3"},{"provider":"OpenRouter","model_id":"google/lyria-3-pro-preview"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Full songs with vocals and lyrics","detail":"Full-length ~2-minute tracks with verses/choruses/bridges, generated vocals and lyrics; 44.1 kHz stereo MP3/WAV.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/music-generation"},{"name":"Image-conditioned music","detail":"Accepts text and image prompts via the Interactions API.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/music-generation"},{"name":"SynthID-watermarked audio","detail":"Latent-diffusion model with SynthID watermarking on outputs.","first":false,"discovered":"launch","source":"https://deepmind.google/models/model-cards/lyria-3-5/"}],"entry":"2026-07-29-google-lyria-3-5","notes":"Launched 2026-07-29 in Google Flow Music (the rebranded ProducerAI); Gemini API GA 2026-09-03 (status Stable, no free tier). Not yet listed on the Vertex/Agent Platform Lyria pages or pricing as of 2026-09-29 (Vertex still offers lyria-3-pro-preview, lyria-3-clip-preview and lyria-002). The lyria-3-clip-preview access line above is the older Lyria 3 Clip, see lyria-3.md; lyria-realtime-exp covers streaming music (lyria-realtime.md). OpenRouter lists only Lyria 3 previews (not 3.5).","verified":"2026-09-29","body_md":"Google's best music generation model (songs with vocals) for apps and creators.\n\n```python\nimport base64\nfrom google import genai\nclient = genai.Client()\nit = client.interactions.create(model=\"lyria-3.5\",\n    input=\"An epic cinematic orchestral piece about a journey home.\")\nopen(\"music.mp3\", \"wb\").write(base64.b64decode(it.output_audio.data))\n```\n\nSources: [music generation guide](https://ai.google.dev/gemini-api/docs/music-generation), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [model card](https://deepmind.google/models/model-cards/lyria-3-5/), [changelog](https://ai.google.dev/gemini-api/docs/changelog).","page_url":"https://postcutoff.com/m/lyria-3-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":600,"you_url":null,"briefings":null},{"id":"gpt-live-transcribe","name":"GPT-Live-Transcribe","org":"OpenAI","family":"GPT Transcribe","released":"2026-07-28","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.017,"unit":"per minute of realtime audio (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.017 per minute","access":[{"provider":"OpenAI API","model_id":"gpt-live-transcribe","endpoint":"v1/realtime/transcription_sessions (Realtime transcription session over WebSocket/WebRTC)","docs":"https://developers.openai.com/api/docs/models/gpt-live-transcribe"}],"capabilities":[{"name":"Low-latency streaming transcription with context hints","detail":"Streams transcript deltas with tunable latency and accepts unstructured context, keyword hints and multiple language hints.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-live-transcribe"},{"name":"Recommended replacement for Whisper streaming use","detail":"Named (with gpt-transcribe) as the replacement for whisper-1 and gpt-4o-(mini-)transcribe(-diarize), which shut down 2027-02-26.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":"2026-07-28-openai-gpt-transcribe-whisper-deprecation","notes":"Released with gpt-transcribe (file transcription, $0.0045/min) on 2026-07-28 per the changelog. Languages and latency figures not published on the docs page.","verified":"2026-09-29","body_md":"Streaming sibling of [gpt-transcribe](gpt-transcribe.md) for captions, voice agents and meeting notes.\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-live-transcribe\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations","page_url":"https://postcutoff.com/m/gpt-live-transcribe/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":605,"you_url":null,"briefings":null},{"id":"gpt-transcribe","name":"GPT-Transcribe","org":"OpenAI","family":"GPT Transcribe","released":"2026-07-28","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.0045,"unit":"per minute of audio","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.0045 per minute","access":[{"provider":"OpenAI API","model_id":"gpt-transcribe","endpoint":"https://api.openai.com/v1/audio/transcriptions","docs":"https://developers.openai.com/api/docs/models/gpt-transcribe"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-transcribe","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Context-guided transcription","detail":"Accepts unstructured context, keyword hints and multiple language hints for domain terms.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-transcribe"},{"name":"Whisper successor","detail":"Replacement for whisper-1 and gpt-4o-(mini-)transcribe (shutdown Feb 26 2027).","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":"2026-07-28-openai-gpt-transcribe-whisper-deprecation","notes":"File and Realtime transcription. Streaming sibling gpt-live-transcribe ($0.017/min). Cheaper than whisper-1 ($0.006/min).","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/audio/transcriptions \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -F model=gpt-transcribe -F file=@audio.mp3\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-transcribe\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-transcribe/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":605,"you_url":null,"briefings":null},{"id":"claude-opus-5","name":"Claude Opus 5","org":"Anthropic","family":"Claude 5","released":"2026-07-24","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-05","pricing":{"input":5,"output":25,"cache_read":0.5,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-opus-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-5","url":"https://openrouter.ai/anthropic/claude-opus-5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Novel problem solving (ARC-AGI 3)","detail":"Anthropic says it scored about 3x as high as competing models on ARC-AGI 3.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-5"},{"name":"Near-Fable coding at half the price","detail":"Launch claim: more than doubles Opus 4.8 on Frontier-Bench and beats Fable 5 on OSWorld 2.0 at about a third of the cost.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-5"},{"name":"Self-built tooling","detail":"In one demo it wrote its own vision pipeline to solve a FreeCAD reconstruction task.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-5"},{"name":"Thinking on by default","detail":"First Opus where omitting the thinking parameter runs adaptive thinking.","first":true,"discovered":"launch","source":"https://platform.claude.com/docs/en/models/opus-5/overview"}],"entry":null,"notes":"Superseded by claude-opus-5-5 (cheaper). Thinking on by default; {type: disabled} allowed only at effort high or below. Fast mode $10/$50 (Claude API only). Retirement not sooner than 2027-07-24.","verified":"2026-09-29","body_md":"Previous Opus. Migrate to `claude-opus-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-5/overview\n- Announcement: https://www.anthropic.com/news/claude-opus-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-5/","events_after":694,"major_after":173,"historic_after":37,"missing_at_launch":74,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-05-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-05-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-05-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-05.md"}},{"id":"midjourney-v8-2","name":"Midjourney V8.2","org":"Midjourney","family":"Midjourney V8","released":"2026-07-24","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Web app","url":"https://www.midjourney.com"},{"provider":"Discord","url":"https://discord.gg/midjourney"}],"capabilities":[{"name":"Instruction-based edit model","detail":"V8.2 edit model (Aug 2026) edits images from plain instructions, takes up to 4 image references (replacing Omni Reference / Character Reference / Retexture) and does inpainting/outpainting.","first":false,"discovered":"later","source":"https://updates.midjourney.com/edit-model-for-v8/"},{"name":"Improved personalization","detail":"V8.2 release focused on aesthetics and personalization profiles that better learn a user's taste from image ratings.","first":false,"discovered":"launch","source":"https://updates.midjourney.com/version-8-2/"},{"name":"Rewritten V8 core with native 2K and better text","detail":"V8 line (alpha 2026-03-17, V8.1 2026-04-14) was rebuilt from scratch: much faster jobs, HD/2K output, better prompt following and in-image text.","first":false,"discovered":"launch","source":"https://updates.midjourney.com/v8-alpha/"}],"entry":null,"notes":"No official public API (web app/Discord only; subscription). V8 alpha 2026-03-17, V8.1 2026-04-14 (default from 2026-06-10), V8.2 2026-07-24 - reportedly now default (not confirmed on an official page). Select with --v 8.2 (syntax per docs; docs page blocked). Pricing not verified.","verified":"2026-09-29","body_md":"Midjourney's latest image model - top-tier aesthetics, personalization (profiles, moodboards, srefs) and an instruction edit model. There is no official public API; use the web app at https://www.midjourney.com (or Discord). Third-party \"Midjourney APIs\" are unofficial and violate ToS.\n\nSources: https://updates.midjourney.com/version-8-2/ , https://updates.midjourney.com/edit-model-for-v8/ , https://updates.midjourney.com/v8-1-is-now-the-default-model/ , https://updates.midjourney.com/v8-alpha/","page_url":"https://postcutoff.com/m/midjourney-v8-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":615,"you_url":null,"briefings":null},{"id":"flux-3","name":"FLUX 3","org":"Black Forest Labs","family":"FLUX 3","released":"2026-07-23","status":"preview","type":"video-gen","modality_in":["text","image","video"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.17,"unit":"per second of video (HD, text/image-to-video); FHD 0.29, QHD 0.40, UHD 0.80, draft 0.06; continuation 0.41/s HD","source":"https://docs.bfl.ai/quick_start/pricing"},"price_line":"$0.17 per second","access":[{"provider":"BFL API","model_id":"flux-3-video","endpoint":"https://api.bfl.ai/v1/flux-3-video","docs":"https://docs.bfl.ai/flux_3/flux3_overview"},{"provider":"Hugging Face (FLUX 3 Action open weights)","url":"https://huggingface.co/black-forest-labs/flux-3-action-base"}],"capabilities":[{"name":"Unified image/video/audio/action model","detail":"Single architecture jointly trained on images, video, audio and robot action prediction; each modality said to strengthen the others.","first":false,"discovered":"launch","source":"https://www.globenewswire.com/news-release/2026/07/23/3332364/0/en/black-forest-labs-unveils-flux-3-a-new-multimodal-frontier-model-for-visual-intelligence.html"},{"name":"Video with native synced audio","detail":"Text/image-to-video up to ~20 s with optional in-sync audio, plus video continuation and video editing (/v1/flux-tools/video-edit-v1).","first":false,"discovered":"launch","source":"https://docs.bfl.ai/flux_3/flux3_overview"},{"name":"FLUX 3 Action for robotics","detail":"Video-prediction engine reused for robot control (FLUX-mimic with mimic robotics, tested by Audi); open-weight Action checkpoints on HF (FLUX Kommunity license).","first":false,"discovered":"launch","source":"https://docs.bfl.ai/flux_3/flux3_action_overview"}],"entry":null,"notes":"Early access at launch (2026-07-23); FLUX 3 Image announced 'in coming weeks' and open FLUX 3 [dev] planned later in 2026 - not verified as released. Action weights: flux-3-action-base/-so101/-droid (HF, 2026-09-22).","verified":"2026-09-29","body_md":"BFL's first video (+audio) model and multimodal \"visual intelligence\" foundation. Async API: submit, then poll the returned `polling_url`.\n\n```bash\ncurl -X POST https://api.bfl.ai/v1/flux-3-video -H \"x-key: $BFL_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"mode\":\"t2v\",\"prompt\":\"a fox running through dawn mist\",\"duration\":5}'\n```\n\nSources: https://docs.bfl.ai/flux_3/flux3_overview , https://docs.bfl.ai/quick_start/pricing , https://www.globenewswire.com/news-release/2026/07/23/3332364/0/en/black-forest-labs-unveils-flux-3-a-new-multimodal-frontier-model-for-visual-intelligence.html","page_url":"https://postcutoff.com/m/flux-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":620,"you_url":null,"briefings":null},{"id":"gemini-3-5-flash-lite","name":"Gemini 3.5 Flash-Lite","org":"Google DeepMind","family":"Gemini 3.5","released":"2026-07-21","status":"current","type":"llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":2.5,"unit":"per 1M tokens (Standard tier)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.30 in, $2.50 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.5-flash-lite","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite"},{"provider":"OpenRouter","model_id":"google/gemini-3.5-flash-lite"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"High-throughput subagent model","detail":"~350 output tokens/s; optimized for subagent tasks and document processing at low cost.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"name":"Strong coding for a Lite tier","detail":"54% on Terminal-Bench 2.1 vs 31% for the previous Flash-Lite.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"entry":null,"notes":"Cheapest current Gemini text model; recommended replacement for 2.5 Flash/Flash-Lite and 3.1 Flash-Lite. Alias gemini-flash-lite-latest may point here (not verified).","verified":"2026-09-29","body_md":"Fast, low-cost multimodal model for classification, extraction, routing and subagents.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Classify: I love this product\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/).","page_url":"https://postcutoff.com/m/gemini-3-5-flash-lite/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":629,"you_url":"https://postcutoff.com/you/gemini-3-5-flash-lite/","briefings":null},{"id":"gemini-3-6-flash","name":"Gemini 3.6 Flash","org":"Google DeepMind","family":"Gemini 3","released":"2026-07-21","status":"legacy","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.75,"output":3.75,"unit":"per 1M tokens (Standard tier, current introductory price through 2026-12-31; $1.50 / $7.50 from 2027-01-01, which was its launch price)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.75 in, $3.75 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.6-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.6-flash:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-6-flash"},{"provider":"OpenRouter","model_id":"google/gemini-3.6-flash"}],"capabilities":[{"name":"Token-efficient agentic coding","detail":"Uses 17% fewer output tokens than 3.5 Flash with better coding/multimodal results and fewer unwanted edits.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"name":"Computer-use agents","detail":"83.0% on OSWorld-Verified (vs 78.4% for 3.5 Flash).","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"}],"entry":null,"notes":"Stable, no shutdown date; superseded by 3.7 and 3.8 Flash. Recommended replacement for gemini-3-flash-preview per deprecations page. Output limit not re-checked (assumed 65,536 like siblings, omitted).","verified":"2026-09-29","body_md":"July 2026 Flash release, launched together with 3.5 Flash-Lite and 3.5 Flash Cyber. Use `gemini-3.8-flash` for new projects.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.6-flash:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [models](https://ai.google.dev/gemini-api/docs/models), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/).","page_url":"https://postcutoff.com/m/gemini-3-6-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":629,"you_url":null,"briefings":null},{"id":"qwen-image-3-0","name":"Qwen-Image-3.0 (Pro)","org":"Alibaba (Qwen)","family":"Qwen-Image","released":"2026-07-21","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.04,"unit":"qwen-image-3.0-pro, per output image at 1K (2K: 0.075; image input 0.003/image), Singapore/International. qwen-image-3.0: 0.03/image at 1K or 2K","source":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},"price_line":"$0.04 per image","access":[{"provider":"Alibaba Cloud Model Studio","model_id":"qwen-image-3.0-pro","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro"},{"provider":"Alibaba Cloud Model Studio","model_id":"qwen-image-3.0","docs":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},{"provider":"Hugging Face (open sibling Qwen-Image-2.1, research license)","url":"https://huggingface.co/Qwen/Qwen-Image-2.1"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"Dense single-pass layouts","detail":"Prompts up to ~4.5K tokens; generates newspapers, storyboards, menus, exam papers and images-within-images in one pass.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro"},{"name":"Tiny, multilingual text rendering","detail":"Legible text down to ~10px, native rendering of 12 languages and multiple fonts, realistic UI simulation (web pages, games, livestreams).","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro"},{"name":"Closed release (break from open Qwen-Image)","detail":"Shipped without weights, benchmarks or model card, unlike earlier open Qwen-Image releases.","first":false,"discovered":"later","source":"https://www.unite.ai/alibaba-launches-qwen-image-3-0-without-benchmarks-or-weights/"}],"entry":null,"notes":"Released 2026-07-21 (invite-only for two weeks, opened to Qwen app users 2026-08-05, per press). Open-weight alternative: Qwen-Image-2.1 (7B DiT, released 2026-09-20 per its GitHub README, qwen-research license; see qwen-image-2-1). API endpoint path not verified here - see docs.","verified":"2026-09-29","body_md":"Text-to-image and image editing aimed at \"useful\" production graphics: infographics, posters, dense text layouts.\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen-image-3-0-pro · https://www.alibabacloud.com/help/en/model-studio/model-pricing · https://www.unite.ai/alibaba-launches-qwen-image-3-0-without-benchmarks-or-weights/","page_url":"https://postcutoff.com/m/qwen-image-3-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":629,"you_url":null,"briefings":null},{"id":"qwen-audio-3-0-tts","name":"Qwen-Audio-3.0-TTS (Flash / Plus)","org":"Alibaba (Qwen / Tongyi Lab)","family":"Qwen-Audio 3.0","released":"2026-07-20","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":27.59,"unit":"USD per 1M characters for the Plus tier at launch (MarkTechPost). Alibaba cut TTS prices about 70% with the 3.1 generation (2026-09-23), so check current pricing","source":"https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/"},"price_line":"$27.59 per 1M characters","access":[{"provider":"Alibaba Cloud Model Studio (Singapore / Beijing)","model_id":"qwen-audio-3.0-tts-flash","endpoint":"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen-tts"},{"provider":"Alibaba Cloud Model Studio","model_id":"qwen-audio-3.0-tts-plus","docs":"https://www.alibabacloud.com/help/en/model-studio/models"}],"capabilities":[{"name":"#1 on Artificial Analysis TTS arena at launch","detail":"Qwen-Audio-3.0-TTS-Plus ranked first on the Artificial Analysis Text-to-Speech leaderboard in July 2026 (Elo ~1,236-1,237, just ahead of Speechify Simba 3.2 at ~1,234). It was later overtaken (Eleven v4 was #1 by late Sept 2026).","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2607.23938"},{"name":"Controllable, robust multilingual synthesis","detail":"12.5 Hz speech tokenizer plus a five-stage LM + flow-matching training recipe; natural-language instructions and inline tags; 16 languages and 20 Chinese dialect regions; one-pass long-form output up to 3 minutes; voice cloning works from noisy or reverberant references.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2607.23938"}],"entry":"2026-07-20-qwen-audio-3-0-tts","notes":"Flash tier targets real-time use (~300 ms first packet, press); Plus targets quality (throughput ~16 chars/s, press). Languages: ar, zh, en, fr, de, id, it, ja, ko, ms, pt, ru, es, tl, th, vi. Companion qwen-audio-3.0-realtime-plus/-flash and qwen-audio-3.0-asr-flash also exist. Superseded by Qwen-Audio-3.1 (2026-09-23), but as of 2026-09-29 the Model Studio catalog still lists qwen-audio-3.0-tts-plus as its TTS model, and no 3.1 TTS id is published in the international docs.","verified":"2026-09-29","body_md":"Alibaba's hosted flagship TTS from July 2026.\n\nSources: https://arxiv.org/abs/2607.23938 · https://www.alibabacloud.com/help/en/model-studio/qwen-tts · https://www.alibabacloud.com/help/en/model-studio/models · https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/","page_url":"https://postcutoff.com/m/qwen-audio-3-0-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":631,"you_url":null,"briefings":null},{"id":"seed-audio-1-0","name":"Seed Audio 1.0","org":"ByteDance","family":"Seed","released":"2026-07-20","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"BytePlus (Seed Speech console)","url":"https://console.byteplus.com/voice/new/setting/activate?projectName=default"}],"capabilities":[{"name":"Unified speech + SFX + ambience generation","detail":"Jointly models voice, sound effects and ambience in one framework for film-grade audio; multi-character dialogue with prompt-level timing control at 100 ms precision; up to 2 min per generation with continuation.","first":false,"discovered":"launch","source":"https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model"},{"name":"20+ languages","detail":"Including zh, en, ja, ko, es, id, de, fr, th, vi; most languages MOS > 4.0 in ByteDance's evaluation.","first":false,"discovered":"launch","source":"https://seed.bytedance.com/en/seedaudio1_0"}],"entry":null,"notes":"API model id and pricing not found. Related ByteDance speech stack: Seed-TTS 2.0 / Doubao TTS 2.0 (Oct 2025), Doubao-Seed-ASR-2.0, Seed LiveInterpret 2.0 (2025-07-24, zh<->en simultaneous interpretation with voice cloning, ~2.5-3 s lag; see entry 2025-07-24-bytedance-seed-liveinterpret-2) on Volcengine/BytePlus. Comparable: Qwen-Audio-3.1-TTS-Next, StepAudio 3 Gen.","verified":null,"body_md":"Sources: https://seed.bytedance.com/en/seedaudio1_0 · https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model","page_url":"https://postcutoff.com/m/seed-audio-1-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":631,"you_url":null,"briefings":null},{"id":"kimi-k3","name":"Kimi K3","org":"Moonshot AI","family":"Kimi K3","released":"2026-07-16","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"kimi-k3 (custom)","context_window":1048576,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":3,"output":15,"cache_read":0.3,"cache_write_5m":3,"cache_write_1h":6,"unit":"per 1M tokens (USD)","source":"https://platform.kimi.ai/docs/pricing/chat"},"price_line":"$3 in, $15 out per 1M tokens","access":[{"provider":"Kimi API (Moonshot)","model_id":"kimi-k3","endpoint":"https://api.moonshot.ai/v1","docs":"https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"},{"provider":"Alibaba Cloud Model Studio","model_id":"kimi-k3","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"OpenRouter","model_id":"moonshotai/kimi-k3","url":"https://openrouter.ai/moonshotai/kimi-k3"},{"provider":"Hugging Face","url":"https://huggingface.co/moonshotai/Kimi-K3"},{"provider":"Web app","url":"https://www.kimi.com"}],"capabilities":[{"name":"First open 3T-class model","detail":"2.8T-parameter MoE (16 of 896 experts active) - Moonshot's claim: the first open model at this scale; weights released after launch (promised by 2026-07-27).","first":true,"discovered":"launch","source":"https://www.kimi.com/blog/kimi-k3"},{"name":"Kimi Delta Attention + Attention Residuals","detail":"Hybrid linear attention (KDA) and AttnRes; ~2.5x the scaling efficiency of K2 per Moonshot.","first":false,"discovered":"launch","source":"https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"},{"name":"Native vision with 1M context","detail":"Native visual understanding (image and video) and a 1,048,576-token window; strong at coding tasks that use screenshots/visual feedback (games, frontend, CAD).","first":false,"discovered":"launch","source":"https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"},{"name":"Always-on thinking with effort control","detail":"Thinking cannot be disabled; reasoning_effort low/high/max (default max).","first":false,"discovered":"launch","source":"https://platform.kimi.ai/docs/guide/kimi-k3-quickstart"}],"entry":null,"notes":"Moonshot flagship. API unlocked after a minimum $1 top-up. Chat Completions, Responses and Anthropic-compatible Messages supported. Max output and knowledge cutoff not verified. Docs moved to platform.kimi.ai (platform.moonshot.ai still serves).","verified":"2026-09-29","body_md":"Moonshot's frontier open model for long-horizon coding, knowledge work and deep reasoning; 1M context, vision.\n\n```bash\ncurl https://api.moonshot.ai/v1/chat/completions \\\n -H \"Authorization: Bearer $MOONSHOT_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"kimi-k3\",\"reasoning_effort\":\"high\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart · https://platform.kimi.ai/docs/pricing/chat · https://platform.kimi.ai/docs/models · https://www.kimi.com/blog/kimi-k3","page_url":"https://postcutoff.com/m/kimi-k3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":636,"you_url":"https://postcutoff.com/you/kimi-k3/","briefings":null},{"id":"minimax-music-3","name":"MiniMax Music 3.0","org":"MiniMax","family":"MiniMax Music","released":"2026-07-16","status":"current","type":"music","modality_in":["text"],"modality_out":["audio"],"open_weights":true,"model_license":"MiniMax-Music3 Community License (commercial use allowed with 'MiniMax-Music3' shown in the UI; separate authorization above US$20M annual revenue)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_song":0.15,"unit":"USD per generation of up to 5 minutes (music-3.0 and music-2.6 API). Paid music APIs closed to NEW users from 2026-08-20; existing paying users keep access.","source":"https://platform.minimax.io/docs/guides/pricing-paygo"},"price_line":"$0.15 per song","access":[{"provider":"MiniMax API (existing paying users only since 2026-08-20)","model_id":"music-3.0","docs":"https://platform.minimax.io/docs/api-reference/music-generation"},{"provider":"Hugging Face","url":"https://huggingface.co/MiniMaxAI/MiniMax-Music3"},{"provider":"GitHub","url":"https://github.com/MiniMax-AI/MiniMax-Music3"},{"provider":"Web app (MiniMax Audio)","url":"https://www.minimax.io/audio"}],"capabilities":[{"name":"Open-weights full songs up to ~5 minutes in one pass","detail":"Composes, arranges, performs and produces a complete song (vocals + arrangement) up to about five minutes from lyrics with section tags and a structured caption; 32 kHz 16-bit stereo WAV.","first":false,"discovered":"launch","source":"https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model"},{"name":"Hierarchical Global/Local LLM with continuous hidden-state synthesis","detail":"8B Global LLM (initialized from Qwen3.5-8B) for long-range structure + 0.6B Local LLM for frame-level acoustics, rendered by a 2.4B flow-matching module and 123M Flow-VAE instead of discrete token decoding.","first":false,"discovered":"launch","source":"https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model"},{"name":"Consumer-GPU inference","detail":"24 GB+ VRAM recommended; runs on 8 GB with CPU offloading; diffusers modular pipeline and ComfyUI support (Comfy-Org/MiniMax-Music-3).","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-Music3"}],"entry":"2026-08-13-minimax-music-3-open-weights","notes":"music-3.0 shipped on the MiniMax API on 2026-07-16 (release notes); open weights published 2026-08-13. Earlier API models: music-2.6 (Apr 2026, covers), music-cover, music-2.5 (Jan 2026), music-2.0 (legacy). On 2026-08-20 MiniMax stopped offering the paid Music and Lyrics Generation APIs to new users and points them to MiniMax Audio or the open model. Demonstrated with English and Mandarin lyrics; no third-party benchmark vs Suno found.","verified":"2026-09-29","body_md":"MiniMax's flagship music model, now the strongest-known open-weights song generator from a major lab. Run locally from Hugging Face (`MiniMaxAI/MiniMax-Music3`) or use the MiniMax Audio web app.\n\nSources: [blog](https://www.minimax.io/blog/minimax-music-3-0-next-generation-open-weights-production-ready-versatile-music-model), [HF model card](https://huggingface.co/MiniMaxAI/MiniMax-Music3), [license](https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE), [models intro](https://platform.minimax.io/docs/guides/models-intro), [release notes](https://platform.minimax.io/docs/release-notes/models), [pricing](https://platform.minimax.io/docs/guides/pricing-paygo).","page_url":"https://postcutoff.com/m/minimax-music-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":636,"you_url":null,"briefings":null},{"id":"mistral-ocr-4-1","name":"Mistral OCR 4.1","org":"Mistral AI","family":"Mistral OCR","released":"2026-07-16","status":"current","type":"multimodal","modality_in":["pdf","image"],"modality_out":["text"],"open_weights":false,"model_license":"mistral-premier","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1000_pages":4,"per_1000_annotated_pages":5,"unit":"USD per 1,000 pages","source":"https://docs.mistral.ai/models/model-cards/ocr-4-1"},"price_line":"$4 per 1,000 pages","access":[{"provider":"Mistral API","model_id":"mistral-ocr-4-1","endpoint":"https://api.mistral.ai/v1/ocr","docs":"https://docs.mistral.ai/models/model-cards/ocr-4-1"}],"capabilities":[{"name":"Paragraph-level bounding boxes with confidence","detail":"Native paragraph-level bbox extraction, structural block labels and block-level confidence scores.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/ocr-4-1"},{"name":"Structured annotations","detail":"Schema-driven document annotation priced separately ($5 / 1,000 annotated pages); batch via /v1/batch.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/ocr-4-1"}],"entry":null,"notes":"Aliases mistral-ocr-4 and mistral-ocr-latest point to 4.1. Powers Mistral Document AI.","verified":"2026-09-29","body_md":"Document OCR to markdown/structured output.\n\n```bash\ncurl https://api.mistral.ai/v1/ocr -H \"Authorization: Bearer $MISTRAL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"mistral-ocr-latest\",\"document\":{\"type\":\"document_url\",\"document_url\":\"https://arxiv.org/pdf/2201.04234\"}}'\n```\n\nSources: https://docs.mistral.ai/models/model-cards/ocr-4-1 , https://docs.mistral.ai/models/overview","page_url":"https://postcutoff.com/m/mistral-ocr-4-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":636,"you_url":"https://postcutoff.com/you/mistral-ocr-4-1/","briefings":null},{"id":"xiaomi-robotics-1","name":"Xiaomi-Robotics-1 (XR-1, 5B)","org":"Xiaomi","family":"Xiaomi-Robotics","released":"2026-07-16","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"XiaomiRobotics/Xiaomi-Robotics-1-5B","url":"https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-1-5B"},{"provider":"GitHub","url":"https://github.com/XiaomiRobotics/Xiaomi-Robotics-1","docs":"https://robotics.xiaomi.com/robot-static-resource/xiaomi-robotics-1/xiaomi-robotics-1.pdf"}],"capabilities":[{"name":"VLA pretrained on 100K+ hours of real trajectories","detail":"Pretrained on 100K+ hours of embodiment-free UMI trajectories across 1,700+ scenarios (per Xiaomi project materials), then post-trained on 10K+ hours of cross-embodiment data, for out-of-the-box mobile manipulation in unseen environments.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2607.15330"},{"name":"Open-weight SOTA on sim benchmarks","detail":"RoboCasa 74.5%, RoboCasa365 57.4%, VLABench 59.1%, RoboDojo 13.93% in the GitHub table (the arXiv abstract cites a 20.07 RoboDojo average score — different metric/version), each ahead of the runner-up per the authors.","first":false,"discovered":"launch","source":"https://github.com/XiaomiRobotics/Xiaomi-Robotics-1"}],"entry":"2026-07-16-xiaomi-robotics-1","notes":"Paper 2026-07-16 (arXiv 2607.15330); weights on HF 2026-07-28; code 2026-08-03. Predecessor: xiaomi-robotics-0 (Feb 2026, arXiv 2602.12684). Companion world model: xiaomi-robotics-u0 (July/Sept 2026). Changelog 2026-09-29: linked the new model files.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/xiaomi-robotics-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":636,"you_url":null,"briefings":null},{"id":"xiaomi-robotics-u0","name":"Xiaomi-Robotics-U0 (38B) / U0-4B","org":"Xiaomi","family":"Xiaomi-Robotics","released":"2026-07-13","status":"current","type":"world-model","modality_in":["text","image"],"modality_out":["image","video","text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"XiaomiRobotics/Xiaomi-Robotics-U0","url":"https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0"},{"provider":"Hugging Face","model_id":"XiaomiRobotics/Xiaomi-Robotics-U0-4B","url":"https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B"},{"provider":"GitHub","url":"https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0"},{"provider":"ModelScope","url":"https://modelscope.cn/collections/XiaomiRobotics/Xiaomi-Robotics-U0"}],"capabilities":[{"name":"Unified embodied synthesis","detail":"One autoregressive model (shared discrete visual tokenizer, next-token objective, initialized from Emu3.5) does text-to-image, image editing, multi-view robot scene generation, embodied transfer (editing scenes while keeping multi-view consistency) and embodied video rollout.","first":true,"discovered":"launch","source":"https://arxiv.org/abs/2607.11643"},{"name":"Data engine for VLAs","detail":"Synthetic data from U0 raised π0.5's out-of-distribution success on hard real-world manipulation tasks from 36.9% to 63.2% (authors).","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2607.11643"},{"name":"FlashAR fast decoding","detail":"Anti-diagonal grouped visual-token decoding plus vLLM batching: 5.44 s per 1024x1024 image on one H20, 82.86x faster than eager AR.","first":false,"discovered":"launch","source":"https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B"}],"entry":"2026-07-13-xiaomi-robotics-u0","notes":"'first' flag is Xiaomi's claim (first model with high-quality multi-view scene generation across multiple robot embodiments). Paper says 38B params; the HF README table says 34B. U0 and U0-FlashAR weights 2026-07-13; U0-4B, U0-Sequence and U0-4B-Sequence weights plus FSDP training code 2026-09-08. U0-Video announced as coming soon. Authors report beating GPT-Image-2.0 in human evals of embodied scene generation/transfer and #1 on World Arena for embodied video. Not an action model: it generates observations/data, not motor commands.","verified":"2026-09-29","body_md":"Xiaomi's open embodied world model / synthetic-data engine, the generation-side companion of its VLAs ([xiaomi-robotics-1](xiaomi-robotics-1.md)).\n\nSources: [arXiv 2607.11643](https://arxiv.org/abs/2607.11643), [HF U0-4B card](https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B), [project page](https://robotics.xiaomi.com/xiaomi-robotics-u0.html).","page_url":"https://postcutoff.com/m/xiaomi-robotics-u0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":645,"you_url":null,"briefings":null},{"id":"cartesia-ink-2","name":"Cartesia Ink-2 (streaming STT)","org":"Cartesia","family":"Ink","released":"2026-07-09","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Cartesia API","model_id":"ink-2","endpoint":"wss://api.cartesia.ai/stt/websocket?model=ink-2","docs":"https://docs.cartesia.ai/build-with-cartesia/stt/latest"},{"provider":"Cartesia API (beta)","model_id":"ink-preview","endpoint":"wss://api.cartesia.ai/stt/websocket"}],"capabilities":[{"name":"Built-in semantic turn detection","detail":"Emits turn.start / turn.update / turn.eager_end / turn.resume / turn.end events so agents need no separate VAD; 89% precision, 93% F1 on endpointing; ~0.1 s time-to-final-transcript.","first":false,"discovered":"launch","source":"https://www.cartesia.ai/blog/introducing-ink-2"},{"name":"#1 streaming WER on Artificial Analysis at launch","detail":"3.4% WER on AA-AgentTalk, ranked #1 on Artificial Analysis's streaming STT leaderboard (company claim, July 2026).","first":false,"discovered":"launch","source":"https://www.cartesia.ai/blog/introducing-ink-2"},{"name":"Keyterm prompting","detail":"Keyterm prompting and configurable turn detection added 2026-08-11.","first":false,"discovered":"later","source":"https://www.cartesia.ai/blog"}],"entry":null,"notes":"Launched English-only (blog 2026-07-09); the current stable `ink-2` snapshot is dated 2026-09-17 and supports English, French, Hindi, Japanese, Spanish. Some press dates an earlier Ink 2 release to May 2026 (unverified). Query params: model, encoding, sample_rate, cartesia_version=2026-08-14; send `finalize` when user stops. Older model: ink-whisper (1 credit/s streaming). Ink-2 credit price not found on pricing page (plans list included STT hours).","verified":"2026-09-29","body_md":"Cartesia's streaming speech-to-text built for voice agents; pairs with Sonic-3.6.\n\nSources: https://docs.cartesia.ai/build-with-cartesia/stt/latest , https://docs.cartesia.ai/api-reference/stt/stt , https://www.cartesia.ai/blog/introducing-ink-2","page_url":"https://postcutoff.com/m/cartesia-ink-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":648,"you_url":null,"briefings":null},{"id":"gpt-5-6-luna","name":"GPT-5.6 Luna","org":"OpenAI","family":"GPT-5.6","released":"2026-07-09","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-02","pricing":{"input":0.2,"cached_input":0.02,"output":1.2,"unit_note":"prompts >272K input tokens billed 2x input / 1.5x output","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.20 in, $1.20 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.6-luna","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.6-luna","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.6-luna","url":"https://openrouter.ai/openai/gpt-5.6-luna"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Budget tier with 1M context","detail":"Fast, low-cost tier with 1.05M context and full reasoning-effort range.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.6-luna"},{"name":"Replacement for gpt-5-nano/mini snapshots","detail":"Named migration target for deprecated small GPT-5 snapshots.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":null,"notes":"Superseded by GPT-6 Luna (half the price) but still available.","verified":"2026-09-29","body_md":"Low-cost high-volume tasks.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.6-luna\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.6-luna\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-6-luna/","events_after":769,"major_after":192,"historic_after":40,"missing_at_launch":116,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-6-luna/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-02-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-02-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-02-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-02.md"}},{"id":"gpt-5-6-sol","name":"GPT-5.6 Sol","org":"OpenAI","family":"GPT-5.6","released":"2026-07-09","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-02","pricing":{"input":4,"cached_input":0.4,"output":20,"unit_note":"prompts >272K input tokens billed 2x input / 1.5x output","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$4 in, $20 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.6-sol","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.6-sol"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.6-sol","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.6-sol","url":"https://openrouter.ai/openai/gpt-5.6-sol"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Named-tier family","detail":"GPT-5.6 introduced the Sol/Terra/Luna tier names (flagship/balanced/fast) replacing pro/mini/nano naming.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/changelog"},{"name":"Max reasoning effort","detail":"Reasoning effort none/low/medium/high/xhigh/max.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.6-sol"},{"name":"Fast mode long context","detail":"Fast mode extended to long-context requests on Aug 5 2026.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/changelog"}],"entry":null,"notes":"GPT-5.6 flagship; the gpt-5.6 alias routes here. Superseded by GPT-6 Sol/Astra but still offered. OpenRouter lists $2/$10, lower than OpenAI list price $4/$20.","verified":"2026-09-29","body_md":"Previous flagship; replacement target for deprecated gpt-5/o3 snapshots.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.6-sol\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.6-sol\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-6-sol/","events_after":769,"major_after":192,"historic_after":40,"missing_at_launch":116,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-6-sol/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-02-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-02-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-02-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-02.md"}},{"id":"gpt-5-6-terra","name":"GPT-5.6 Terra","org":"OpenAI","family":"GPT-5.6","released":"2026-07-09","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2026-02","pricing":{"input":2,"cached_input":0.2,"output":12,"unit_note":"prompts >272K input tokens billed 2x input / 1.5x output","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2 in, $12 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.6-terra","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.6-terra"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.6-terra","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.6-terra","url":"https://openrouter.ai/openai/gpt-5.6-terra"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Balanced tier","detail":"Mid tier at $2/$12 per 1M, well below GPT-5.5 ($5/$30), with max reasoning effort.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/pricing"},{"name":"Official migration target","detail":"Named replacement for many deprecated legacy snapshots (gpt-3.5, gpt-4 variants, o-series).","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":null,"notes":"Balanced GPT-5.6 model; no GPT-6 Terra counterpart as of 2026-09-29. Single snapshot gpt-5.6-terra.","verified":"2026-09-29","body_md":"Everyday work at mid-range cost.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.6-terra\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.6-terra\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-6-terra/","events_after":769,"major_after":192,"historic_after":40,"missing_at_launch":116,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-6-terra/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-02-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-02-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-02-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-02.md"}},{"id":"gpt-live-1","name":"GPT-Live 1","org":"OpenAI","family":"GPT Live","released":"2026-07-08","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":"2025-07","pricing":{"per_minute":0.05,"unit":"per minute of voice session (billed per second); backend model and tools billed separately","source":"https://developers.openai.com/api/docs/models/gpt-live-1"},"price_line":"$0.05 per minute","access":[{"provider":"OpenAI API","model_id":"gpt-live-1","endpoint":"https://api.openai.com/v1/live/sessions","docs":"https://developers.openai.com/api/docs/models/gpt-live-1"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Full-duplex voice","detail":"Listens and speaks at the same time, delegating reasoning and tool use to a backend agent model.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-live-1"},{"name":"New Live API","detail":"Served on a dedicated v1/live/sessions endpoint rather than Realtime.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-live-1"},{"name":"Replaced turn-based Advanced Voice Mode in ChatGPT","detail":"Since 2026-07-08 GPT-Live-1 (paid tiers) and GPT-Live-1 mini (default, all users) power ChatGPT Voice, with backchannels ('mhmm') and background hand-off of hard questions to GPT-5.5.","first":false,"discovered":"launch","source":"https://techcrunch.com/2026/07/08/openai-releases-new-voice-models-for-more-natural-live-conversations/"}],"entry":"2026-07-08-openai-gpt-live-chatgpt-voice","notes":"Launched in ChatGPT 2026-07-08 (GPT-Live-1 for Go/Plus/Pro, GPT-Live-1 mini default for Free); ChatGPT desktop (macOS/Windows) ~2026-07-23; API GA 2026-09-10 per changelog (earlier preview around 2026-07-31). gpt-live-1-mini is ChatGPT-only: not in the API models catalog and developers.openai.com/api/docs/models/gpt-live-1-mini returns 404 (checked 2026-09-29). ChatGPT Voice limits (Unite.AI, 2026-09-23): Free limited mini, Go 3 h mini, Plus 3 h GPT-Live-1, Pro $100 15 h, Pro $200 unlimited; Enterprise/Edu 1.25 credits/min or $0.05/min. Since 2026-09-23 Voice can use plugins/connected apps and runs inside ChatGPT Work. Knowledge cutoff 2025-07-31. Concurrency 25-500 sessions by tier. No image/video input. Not listed on Azure or OpenRouter. OpenAI's launch post returned 403 to our fetcher; ChatGPT facts from TechCrunch.","verified":"2026-09-29","body_md":"Natural full-duplex voice assistants that hand hard work to a text model (e.g. gpt-6-sol).\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-live-1\n- https://developers.openai.com/api/docs/changelog\n- https://www.unite.ai/openai-brings-plugins-to-live-voice-and-voice-to-work-in-chatgpt/\n- https://openai.com/index/introducing-gpt-live/\n- https://techcrunch.com/2026/07/08/openai-releases-new-voice-models-for-more-natural-live-conversations/\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-live-1/","events_after":853,"major_after":216,"historic_after":43,"missing_at_launch":196,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-07-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-07-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-07-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-07.md"}},{"id":"robostral-navigate","name":"Robostral Navigate","org":"Mistral AI","family":"Robostral","released":"2026-07-08","status":"preview","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Mistral AI (contact sales / partners; no public API id or weights found)","url":"https://mistral.ai/news/robostral-navigate/","docs":"https://arxiv.org/abs/2607.20785"}],"capabilities":[{"name":"Single-RGB-camera vision-language navigation","detail":"8B model navigates buildings from one RGB camera plus language instructions (no LiDAR/depth); R2R-CE success 79.4% val-seen, 76.6% val-unseen (+9.7 pts over best single-camera method, +4.5 over depth/multi-camera systems).","first":false,"discovered":"launch","source":"https://mistral.ai/news/robostral-navigate/"},{"name":"Sim-only training, embodiment-agnostic","detail":"Trained in simulation (~2.4M trajectories across 350k scenes per Mistral's page), with prefix caching (22x fewer training tokens) and online RL (CISPO, +3.2 pts); works on wheeled, legged and flying robots.","first":false,"discovered":"launch","source":"https://mistral.ai/news/robostral-navigate/"}],"entry":"2026-07-08-mistral-robostral-navigate","notes":"Mistral's first robotics model; built in-house without an existing open VLM. Outputs navigation actions. Access appears to be via Mistral's team ('talk with our team'); status set to preview.","verified":"2026-09-29","body_md":"Sources: [Mistral: Robostral Navigate](https://mistral.ai/news/robostral-navigate/), [tech report](https://arxiv.org/abs/2607.20785).","page_url":"https://postcutoff.com/m/robostral-navigate/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":653,"you_url":null,"briefings":null},{"id":"assemblyai-universal-3-5-pro","name":"AssemblyAI Universal-3.5 Pro (async)","org":"AssemblyAI","family":"Universal-3","released":"2026-07-07","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.21,"unit":"per audio hour (pre-recorded); keyterms prompting +$0.05/hr, diarization +$0.02/hr, Medical Mode +$0.15/hr","source":"https://www.assemblyai.com/pricing"},"price_line":"$0.21 per hour","access":[{"provider":"AssemblyAI API","model_id":"universal-3-pro","endpoint":"https://api.assemblyai.com/v2/transcript","docs":"https://www.assemblyai.com/docs/getting-started/universal-3-5-pro"},{"provider":"OpenRouter","url":"https://openrouter.ai/assemblyai/universal-3-5-pro"},{"provider":"AssemblyAI Dictation API","model_id":"(Universal-3.5 Pro + LLM cleanup)","endpoint":"https://dictation.assemblyai.com/v1/transcribe/live","docs":"https://www.assemblyai.com/blog/dictation-api"}],"capabilities":[{"name":"Promptable speech language model","detail":"Universal-3 Pro (Feb 2026) introduced plain-language prompts controlling transcription (disfluencies, multilingual handling, PII, formatting); 3.5 Pro focuses on entities, rare words and domain terms with an LLM-based decoder.","first":false,"discovered":"launch","source":"https://www.assemblyai.com/blog/introducing-universal-3-pro"},{"name":"Dictation API (polished text from short utterances)","detail":"Launched 2026-09-15: up to 5 s audio per request (chunked upload), removes fillers, resolves self-corrections and fixes name spellings via `llm_instruction`, `keyterms_prompt` and `stt_prompt`; 0.36 s average response, 3.87% WER on short-form audio (vendor-cited), 19 languages, $0.62/hour all-in. Open-source MIT macOS demo app 'Blurt'.","first":false,"discovered":"later","source":"https://www.assemblyai.com/blog/dictation-api"},{"name":"Medical Mode","detail":"`domain: medical-v1` for EN/ES/DE/FR clinical vocabulary; replaces deprecated Slam-1.","first":false,"discovered":"later","source":"https://www.assemblyai.com/llms/models.md"}],"entry":null,"notes":"API id stays `universal-3-pro` (pass in `speech_models`, plural; singular `speech_model` is deprecated). 18 languages; use universal-2 ($0.15/hr, 99+ languages) for broad coverage and legacy features (auto_chapters/summarization fail on 3.5 Pro). Added to OpenRouter 2026-09-22. Launch date 2026-07-07 from AssemblyAI releases collection via search (not opened directly). Streaming sibling: assemblyai-universal-3-6-pro-realtime. Related AssemblyAI products: Voice Agent API (GA April 2026, $4.50/hr all-in) and LLM Gateway (OpenAI-compatible multi-provider LLM API that replaced LeMUR; migration guide at assemblyai.com/docs/llm-gateway/migration-from-lemur; exact rename date not verified).","verified":"2026-09-29","body_md":"Sources: https://www.assemblyai.com/llms/models.md , https://www.assemblyai.com/pricing","page_url":"https://postcutoff.com/m/assemblyai-universal-3-5-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":657,"you_url":null,"briefings":null},{"id":"speechify-simba-3-2","name":"Speechify Simba 3.2","org":"Speechify (SpeechifyAI)","family":"Simba 3","released":"2026-07-07","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":10,"unit":"USD per 1M characters (entry tier); $6 at scale tier","source":"https://www.prweb.com/releases/speechifys-simba-3-2-ranks-1-on-independent-artificial-analysis-tts-leaderboard-worlds-best-real-time-voice-model-above-elevenlabs-openai-google-deepmind--others-302819731.html"},"price_line":"$10 per 1M characters","access":[{"provider":"SpeechifyAI API","endpoint":"https://api.speechify.ai/v1/audio/speech (also /v1/audio/stream)","docs":"https://docs.speechify.ai/build/guides/concepts/models"},{"provider":"Web","url":"https://speechify.ai/models"}],"capabilities":[{"name":"Briefly #1 on Artificial Analysis Speech Arena at a low price","detail":"Press release 2026-07-07 claimed #1 on the AA TTS leaderboard; a week later Qwen-Audio-3.0-TTS-Plus overtook it (1,236 vs 1,234 Elo). On 2026-09-29 it was #7 (Elo 1239). Speechify called it the cheapest model in the top ten ($10/$6 per 1M chars).","first":false,"discovered":"launch","source":"https://artificialanalysis.ai/text-to-speech/leaderboard"},{"name":"Streaming-native, low TTFB","detail":"Streaming-native Simba 3 model; <100 ms first byte claimed; emotional control, SSML prosody, instant voice cloning; 30+ locales with mixed-language input. Recommended model for English integrations.","first":false,"discovered":"launch","source":"https://speechify.ai/blog/simba-3-2-streaming-model"}],"entry":null,"notes":"Exact API model id string not verified (docs page 'SpeechifyAI Build TTS Models: Simba 3.2, 3.0, Multilingual, and English'). AA measured ~30.2 chars/s generation speed (the-decoder, Jul 2026). Quotes: Luke Oliff, Tyler Weitzman in the press release.","verified":"2026-09-29","body_md":"Sources: https://speechify.ai/blog/simba-3-2-and-the-models-endpoint , https://docs.speechify.ai/build/guides/concepts/models , PRWeb release (2026-07-07), https://artificialanalysis.ai/text-to-speech/leaderboard","page_url":"https://postcutoff.com/m/speechify-simba-3-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":657,"you_url":null,"briefings":null},{"id":"gpt-realtime-2-1","name":"GPT-Realtime-2.1","org":"OpenAI","family":"GPT Realtime","released":"2026-07-06","status":"current","type":"audio/speech","modality_in":["text","audio","image"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":32000,"knowledge_cutoff":"2024-09","pricing":{"text_input":4,"text_output":24,"audio_input":32,"audio_output":64,"cached_input":0.4,"image_input":5,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$4 text in, $24 text out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-realtime-2.1","endpoint":"wss://api.openai.com/v1/realtime","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-2.1"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-realtime-2.1","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Reasoning in realtime voice","detail":"Configurable reasoning effort in a speech-to-speech model (at a latency cost).","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-2.1"},{"name":"Robust turn-taking","detail":"Improved alphanumeric recognition, silence/noise handling and interruption behavior.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-2.1"}],"entry":null,"notes":"Realtime API only. Successor to gpt-realtime-2 (2026-05-07, same prices, see gpt-realtime-2.md). Mini variant gpt-realtime-2.1-mini (audio $10/$20, text $0.60/$2.40). Replaces gpt-realtime / gpt-4o-realtime (shutdown Jan 20 2027). Azure version 2026-07-07.","verified":"2026-09-29","body_md":"Low-latency voice agents over WebRTC/WebSocket (`/v1/realtime?model=gpt-realtime-2.1`).\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-realtime-2.1\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-realtime-2-1/","events_after":897,"major_after":238,"historic_after":51,"missing_at_launch":237,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-09.md"}},{"id":"leanstral-1-5","name":"Leanstral 1.5","org":"Mistral AI","family":"Leanstral","released":"2026-07-02","status":"current","type":"code","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0,"output":0,"unit":"free API endpoint (per Mistral's announcement)","source":"https://mistral.ai/news/leanstral-1-5"},"price_line":"Free","access":[{"provider":"Mistral API","model_id":"leanstral-1-5","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://mistral.ai/news/leanstral-1-5"},{"provider":"Hugging Face","url":"https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B"}],"capabilities":[{"name":"Lean 4 proof agent","detail":"Specialised for Lean 4 theorem proving in real repositories: miniF2F 100%, PutnamBench 587/672, FATE-H 87%, FATE-X 34% (Mistral-reported).","first":false,"discovered":"launch","source":"https://mistral.ai/news/leanstral-1-5"},{"name":"Formal code verification","detail":"Used as a verification agent over 57 open-source repos, it found 5 previously unreported bugs (e.g. an integer overflow in datrs/varinteger).","first":false,"discovered":"launch","source":"https://mistral.ai/news/leanstral-1-5"}],"entry":"2026-07-02-mistral-leanstral-1-5","notes":"119B MoE (128 experts, 4 active, ~6B active params). Update of Leanstral (Leanstral-2603, March 16, 2026). 256K context, ≤200K recommended.","verified":"2026-10-01","body_md":"Open-weights Lean 4 prover and verification agent from Mistral. Weights are on Hugging Face (Apache 2.0), and Mistral's API offers a free endpoint (`leanstral-1-5`).","page_url":"https://postcutoff.com/m/leanstral-1-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":660,"you_url":null,"briefings":null},{"id":"mureka-9-5","name":"Mureka V9.5 (and O3)","org":"Kunlun Tech (Skywork AI)","family":"Mureka","released":"2026-07","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Mureka API","model_id":"mureka-9.5","docs":"https://platform.mureka.ai/docs/"},{"provider":"Web app","url":"https://www.mureka.ai"}],"capabilities":[{"name":"MusiCoT (music chain-of-thought) planning","detail":"Mureka's line plans song structure, sections and intent before generating audio (MusiCoT); Mureka O1 (2025-07-29) was billed as the first 'thinking' music reasoning model, followed by O2 (2025-12-09) and O3 'reflective reasoning' with V9.5.","first":false,"discovered":"launch","source":"https://www.prnewswire.com/news-releases/kunlun-tech-launches-the-worlds-first-music-reasoning-large-model-mureka-o1-leading-the-global-ai-music-revolution-302411665.html"},{"name":"MuCo creation agent","detail":"Agent that manages a song as a version-controlled project instead of one-shot generation (per Pandaily/Variety coverage of V9.5).","first":false,"discovered":"launch","source":"https://pandaily.com/mureka-v9-5-ai-music-kunlun-tech-jul2026"},{"name":"Fine-tuning API and vocal cloning","detail":"API offers song/instrumental/lyrics generation, song extension, stem separation, transcription, vocal cloning and custom-model fine-tuning on 200+ consistent tracks.","first":false,"discovered":"launch","source":"https://platform.mureka.ai/docs/"}],"entry":null,"notes":"Release history per the API changelog (https://platform.mureka.ai/docs/en/changelog.html): mureka-7 + mureka-o1 2025-07-29; mureka-7.5 2025-09-25; mureka-7.6 + mureka-o2 2025-12-09; mureka-8 2026-03-02 (consumer Mureka V8 announced 2026-01-28, claimed to surpass Suno in melody, vocals, arrangement and emotion; cited as a baseline in Tencent's SongGeneration 2 paper); mureka-9 2026-04-09; enhanced mureka-9.5 2026-08-28. V9.5 was shown around WAIC (late July 2026) and formally announced 2026-08-31 (GlobeNewswire) with internal-test figures: 61.0% of lead vocals rated convincing, 97.0% prompt following, 95.7% genre match. Exact consumer launch day and API pricing not verified; training-data provenance undisclosed. Kunlun Tech's music models are developed under its Skywork AI unit.","verified":"2026-09-29","body_md":"Kunlun Tech's commercial song generator family, a leading Chinese Suno competitor with a public developer API (unlike Suno). Use via https://www.mureka.ai or the API (model id `mureka-9.5`; older `mureka-9`, `mureka-8`, `mureka-7.5`, `mureka-o2` also referenced in docs).\n\nSources: https://platform.mureka.ai/docs/en/changelog.html , https://www.globenewswire.com/news-release/2026/08/31/3353336/0/en/mureka-introduces-next-generation-ai-music-model-v9-5-advancing-toward-more-human-like-song-creation.html , https://en.youth.cn/RightNow/202601/t20260130_16489615.htm , https://variety.com/2026/shopping/news/ai-music-generator-mureka-v9-5-o3-models-launch-details-1236839140/","page_url":"https://postcutoff.com/m/mureka-9-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":595,"you_url":null,"briefings":null},{"id":"claude-sonnet-5","name":"Claude Sonnet 5","org":"Anthropic","family":"Claude 5","released":"2026-06-30","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-01","pricing":{"input":2,"output":10,"cache_read":0.2,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$2 in, $10 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-sonnet-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/sonnet-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-sonnet-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-sonnet-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-sonnet-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-sonnet-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-sonnet-5","url":"https://openrouter.ai/anthropic/claude-sonnet-5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Opus 4.8-level quality at Sonnet price","detail":"Anthropic positioned it at parity with Opus 4.8 on many tasks at $2/$10 per MTok.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-5"},{"name":"Self-verification","detail":"Early testers reported it checks its own output without prompting and finishes multi-step workflows where earlier Sonnets stopped short.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-5"},{"name":"Cyber safeguards on by default","detail":"It launched with deliberately reduced exploit-development capability and with cyber safeguards enabled.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-5"}],"entry":null,"notes":"Superseded by claude-sonnet-5-5. $2/$10 introductory price became permanent (planned increase to $3/$15 cancelled). Retirement not sooner than 2027-06-30.","verified":"2026-09-29","body_md":"Previous Sonnet. Migrate to `claude-sonnet-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-sonnet-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/sonnet-5/overview\n- Announcement: https://www.anthropic.com/news/claude-sonnet-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-sonnet-5/","events_after":793,"major_after":197,"historic_after":40,"missing_at_launch":125,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-sonnet-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-01.md"}},{"id":"gemini-3-1-flash-lite-image","name":"Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image)","org":"Google DeepMind","family":"Gemini Image","released":"2026-06-30","status":"current","type":"image-gen","modality_in":["text","image","video"],"modality_out":["image","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.25,"output":1.5,"per_image_1k":0.0336,"unit":"per 1M tokens text/image/video input and text output; image output $30 per 1M tokens (~$0.0336 per 1K image)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.25 in, $1.50 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.1-flash-lite-image","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite-image:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite-image"},{"provider":"OpenRouter","model_id":"google/gemini-3.1-flash-lite-image"}],"capabilities":[{"name":"Lowest-cost Gemini image model","detail":"About half the per-image price of Nano Banana 2 (~$0.034 per 1K image).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/pricing"},{"name":"Video-as-input image generation","detail":"Accepts text, image and video inputs for image generation/editing.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/pricing"}],"entry":null,"notes":"GA in Gemini API 2026-06-30 per changelog. Token limits not verified on docs (OpenRouter lists 65,536 context).","verified":"2026-09-29","body_md":"High-volume, low-cost image generation/editing.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite-image:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Product shot of a red sneaker on white\"}]}]}'\n```\n\nSources: [models](https://ai.google.dev/gemini-api/docs/models), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [changelog](https://ai.google.dev/gemini-api/docs/changelog).","page_url":"https://postcutoff.com/m/gemini-3-1-flash-lite-image/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":665,"you_url":null,"briefings":null},{"id":"speechmatics-melia-1","name":"Speechmatics Melia 1 (multilingual STT)","org":"Speechmatics","family":"Speechmatics STT","released":"2026-06-17","status":"preview","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.129,"unit":"USD per audio hour (starting price), 10 hours/month free","source":"https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model"},"price_line":"$0.129 per hour","access":[{"provider":"Speechmatics Batch API","model_id":"melia-1","endpoint":"batch jobs with \"model\": \"melia-1\" and \"language\": \"multi\" (EU1, US1)","docs":"https://docs.speechmatics.com/speech-to-text/models"}],"capabilities":[{"name":"Code-switching across 55+ languages without language selection","detail":"Transcribes audio that switches languages mid-conversation with no language pre-selection; Speechmatics reports it beats Deepgram and Microsoft on 91% and AssemblyAI on 77% of FLEURS languages, and 5% lower WER than its Standard model on noisy monolingual audio.","first":false,"discovered":"launch","source":"https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model"}],"entry":null,"notes":"Launched 2026-06-17 as a production preview (docs: early access), batch only; runs alongside the Standard and Enhanced models. Benchmarks are vendor-reported.","verified":"2026-09-29","body_md":"Sources: https://www.speechmatics.com/company/articles-and-news/introducing-melia-multilingual-speech-to-text-model , https://docs.speechmatics.com/speech-to-text/models","page_url":"https://postcutoff.com/m/speechmatics-melia-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":673,"you_url":null,"briefings":null},{"id":"soniox-stt-v5","name":"Soniox v5 (Async and Real-Time STT)","org":"Soniox","family":"Soniox STT","released":"2026-06-11","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1.5,"output":3.5,"unit":"USD per 1M tokens for async (audio in $1.50, text in/out $3.50; ~$0.10 per audio hour). Real-time: $2.00 audio in, $4.00 text in/out (~$0.12/hour). 1 hour of audio ≈ 30,000 input tokens.","source":"https://soniox.com/pricing"},"price_line":"$1.50 in, $3.50 out per 1M tokens","access":[{"provider":"Soniox API (async / file)","model_id":"stt-async-v5","docs":"https://soniox.com/docs/stt/models"},{"provider":"Soniox API (real-time streaming)","model_id":"stt-rt-v5","docs":"https://soniox.com/docs/stt/models"},{"provider":"Web app","url":"https://soniox.com"}],"capabilities":[{"name":"One multilingual model for 60+ languages with speaker separation","detail":"Soniox claims native-speaker accuracy across 60+ languages in a single model, re-engineered speaker diarization, spoken-language ID, context injection and precise alphanumerics (IDs, emails, codes).","first":false,"discovered":"launch","source":"https://soniox.com/blog/soniox-v5-async"},{"name":"Real-time translation and semantic endpointing","detail":"stt-rt-v5 transcribes and translates live across ~3,600 language pairs, with a tunable `endpoint_sensitivity` semantic endpointing parameter for voice agents.","first":false,"discovered":"launch","source":"https://soniox.com/blog/soniox-v5-real-time"}],"entry":null,"notes":"stt-async-v5 released 2026-06-11, stt-rt-v5 on 2026-06-16. The v4 ids (stt-async-v4 from 2026-01-29, stt-rt-v4 from 2026-02-05) were retired 2026-06-30 and are now aliases routing to v5. Launch posts give no WER numbers; Soniox publishes its own comparisons at soniox.com/benchmarks (vendor-run). Sibling TTS: soniox-tts-v2.","verified":"2026-09-29","body_md":"Soniox's speech-to-text generation for 2026: a single multilingual model family for files (async) and streaming (real-time).\n\nSources: https://soniox.com/blog/soniox-v5-async , https://soniox.com/blog/soniox-v5-real-time , https://soniox.com/docs/stt/models , https://x.com/soniox_ai/status/2065083564027257182","page_url":"https://postcutoff.com/m/soniox-stt-v5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":679,"you_url":null,"briefings":null},{"id":"claude-fable-5","name":"Claude Fable 5","org":"Anthropic","family":"Claude 5","released":"2026-06-09","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-01","pricing":{"input":10,"output":50,"cache_read":1,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$10 in, $50 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-fable-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/fable-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-fable-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-fable-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-fable-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-fable-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-fable-5","url":"https://openrouter.ai/anthropic/claude-fable-5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Strongest cybersecurity capabilities (Mythos 5)","detail":"Anthropic called the Fable 5 / Mythos 5 generation the 'strongest cybersecurity capabilities of any model in the world'. Mythos 5 runs without safety classifiers for Glasswing defenders.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"name":"Rebuild web apps from screenshots","detail":"Anthropic claims it was the first model to rebuild a web app's source code from screenshots alone. It also completed Pokemon FireRed using vision only.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"name":"Massive code migrations","detail":"Stripe reported a 50-million-line migration done in one day instead of about two months.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"name":"Novel scientific hypotheses","detail":"In blind comparisons, scientists preferred its molecular-biology hypotheses about 80% of the time over Opus-class models.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"name":"Refusal stop reason with fallbacks","detail":"Safety classifiers can decline a request with stop_reason 'refusal'. A server-side fallbacks parameter retries on another Claude model.","first":true,"discovered":"launch","source":"https://platform.claude.com/docs/en/models/fable-5/overview"}],"entry":null,"notes":"Superseded by claude-fable-5-1 (same price, cheaper cache reads). Still served; retirement not sooner than 2027-06-09. Sibling claude-mythos-5 (Project Glasswing only, no safety classifiers).","verified":"2026-09-29","body_md":"Previous top-tier model. Prefer `claude-fable-5-1` for new work.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-fable-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/fable-5/overview\n- Announcement: https://www.anthropic.com/news/claude-fable-5-mythos-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-fable-5/","events_after":793,"major_after":197,"historic_after":40,"missing_at_launch":109,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-fable-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-01.md"}},{"id":"gemini-3-5-live-translate","name":"Gemini 3.5 Live Translate","org":"Google DeepMind","family":"Gemini 3.5","released":"2026-06-09","status":"preview","type":"audio/speech","modality_in":["audio"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":65536,"knowledge_cutoff":null,"pricing":{"audio_input":3.5,"audio_output":21,"unit":"per 1M tokens (USD); ~$0.0053/min audio in, ~$0.0315/min audio out","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$3.50 audio in, $21 audio out per 1M tokens","access":[{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-3.5-live-translate-preview","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-live-translate-preview"},{"provider":"Google AI Studio","url":"https://aistudio.google.com/live?model=gemini-3.5-live-translate-preview"},{"provider":"Google Translate app / Google Meet","url":"https://translate.google.com"}],"capabilities":[{"name":"Continuous speech-to-speech translation preserving the speaker's voice","detail":"Audio-to-audio (no ASR-MT-TTS cascade), generating speech continuously a few seconds behind the speaker while keeping intonation, pacing and pitch.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/"},{"name":"70+ languages, 2,000+ pairs","detail":"Auto-detects 70+ languages and supports 2,000+ language combinations in one meeting; expands Google Meet live translation from 5 to 70+ languages.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/"}],"entry":"2026-06-09-gemini-3-5-live-translate","notes":"Public preview in the Live API/AI Studio from 2026-06-09; Meet private preview; Google Translate on Android/iOS (incl. headphone 'listening mode'). Outputs SynthID-watermarked. No function calling, thinking or caching. Model card (read 2026-09-29): https://deepmind.google/models/model-cards/gemini-3-5-audio/ - covers Gemini 3.5 Live Translate, Transcribe and Transcribe Live (card dated 2026-08-26); no numeric evals in the card itself; knowledge cutoff January 2025; did not reach any Tracked or Critical Capability Levels under the Frontier Safety Framework. Card-listed Live Translate limitations: inconsistent voices, language detection struggles with non-native accents and rapid switching, imperfect background-noise handling, occasional audio artifacts. OpenAI's rival gpt-realtime-translate launched a month earlier (2026-05-07).","verified":"2026-09-29","body_md":"Sources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-live-translate-preview), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/).","page_url":"https://postcutoff.com/m/gemini-3-5-live-translate/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":681,"you_url":null,"briefings":null},{"id":"luma-ray-3-2","name":"Luma Ray3.2","org":"Luma AI","family":"Ray3","released":"2026-06-09","status":"current","type":"video-gen","modality_in":["text","image","video"],"modality_out":["video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Luma API","model_id":"ray-3.2","endpoint":"https://agents.lumalabs.ai/v1/generations","docs":"https://docs.agents.lumalabs.ai/"},{"provider":"Web app (Dream Machine)","url":"https://app.lumalabs.ai"}],"capabilities":[{"name":"Multi-keyframe direction","detail":"Up to 16 keyframes inside a single clip for frame-level control of how action evolves.","first":false,"discovered":"launch","source":"https://lumalabs.ai/news/introducing-ray-3-2"},{"name":"Native HDR with 16-bit EXR export","detail":"Generates native HDR video with 16-bit EXR export for pro post-production; up to 20 s at 1080p.","first":false,"discovered":"launch","source":"https://lumalabs.ai/news/introducing-ray-3-2"},{"name":"Multi-face performance tracking and reframe","detail":"Performance tracking for up to 8 faces and an improved reframe tool; full Ray control surface exposed via API for the first time.","first":false,"discovered":"launch","source":"https://lumalabs.ai/news/introducing-ray-3-2"}],"entry":null,"notes":"Successor of Ray3 / Ray3 Modify / Ray3.14. Same API also serves image models uni-1 and uni-1-max (UNI-1.1). Credit-based API pricing (https://lumalabs.ai/pricing) - per-second price not verified.","verified":"2026-09-29","body_md":"Luma's production video model (generation, editing, reframing) behind Dream Machine.\n\n```bash\ncurl -X POST https://agents.lumalabs.ai/v1/generations -H \"Authorization: Bearer $LUMA_AGENTS_API_KEY\" \\\n  -H \"Content-Type: application/json\" -d '{\"model\":\"ray-3.2\",\"prompt\":\"a hot air balloon over salt flats at sunrise\"}'\n# poll GET https://agents.lumalabs.ai/v1/generations/{generation_id}\n```\n\nSources: https://lumalabs.ai/news/introducing-ray-3-2 , https://docs.agents.lumalabs.ai/ , https://lumalabs.ai/llm-info","page_url":"https://postcutoff.com/m/luma-ray-3-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":681,"you_url":null,"briefings":null},{"id":"higgs-audio-v3","name":"Boson AI Higgs Audio v3 (Higgs TTS 3 4B / Higgs STT 3)","org":"Boson AI","family":"Higgs Audio","released":"2026-06-04","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio","text"],"open_weights":true,"model_license":"Boson Higgs TTS 3 Research and Non-Commercial License (TTS weights; Creator Use Grant for attributed monetized content)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"bosonai/higgs-audio-v3-tts-4b","url":"https://huggingface.co/bosonai/higgs-audio-v3-tts-4b"},{"provider":"GitHub","url":"https://github.com/boson-ai/higgs-audio"},{"provider":"SGLang-Omni","docs":"https://www.lmsys.org/blog/2026-06-04-higgs-audio-v3-tts/"}],"capabilities":[{"name":"102-language expressive TTS with zero-shot cloning","detail":"~4B AR decoder (24 kHz, 8 codebooks); 85 languages at production quality (WER/CER <5%), 17 usable; inline control of emotion, style, prosody, pauses and sound effects; 8K-token context; sub-second TTFA streaming.","first":false,"discovered":"launch","source":"https://huggingface.co/bosonai/higgs-audio-v3-tts-4b"},{"name":"Higgs STT 3 (API)","detail":"Speech-to-text model (2026-03-18) for 94 languages; 1.55% WER on LibriSpeech test-clean vs 2.10% for Whisper-large-v3 (company figures). No open weights found.","first":false,"discovered":"launch","source":"https://www.boson.ai/blog/higgs-audio-v3-stt"}],"entry":null,"notes":"TTS weights non-commercial; production/hosted use needs a Boson commercial license or the Boson API (pricing not found). Also mirrored as bosonai/higgs-tts-3-4b. Predecessor Higgs Audio v2 (2025, Apache-2.0-style) on the same GitHub.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/bosonai/higgs-audio-v3-tts-4b , https://www.boson.ai/blog/higgs-audio-v3-tts , https://www.boson.ai/blog/higgs-audio-v3-stt","page_url":"https://postcutoff.com/m/higgs-audio-v3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":688,"you_url":null,"briefings":null},{"id":"nemotron-3-ultra","name":"NVIDIA Nemotron 3 Ultra (550B-A55B)","org":"NVIDIA","family":"Nemotron 3","released":"2026-06-04","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"openmdw-1.1","context_window":1000000,"max_output":null,"knowledge_cutoff":"2025-09","pricing":{"input":0.6,"output":2.4,"unit":"per 1M tokens (USD) on OpenRouter (262K context there); NVIDIA hosted pricing not verified","source":"https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b"},"price_line":"$0.60 in, $2.40 out per 1M tokens","access":[{"provider":"NVIDIA API (build.nvidia.com)","model_id":"nvidia/nemotron-3-ultra-550b-a55b","endpoint":"https://integrate.api.nvidia.com/v1/chat/completions","docs":"https://build.nvidia.com"},{"provider":"OpenRouter","model_id":"nvidia/nemotron-3-ultra-550b-a55b","url":"https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b"},{"provider":"Hugging Face","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"}],"capabilities":[{"name":"Hybrid Mamba-2 / LatentMoE at frontier scale","detail":"550B total / 55B active; interleaved Mamba-2 and LatentMoE layers with select attention, plus multi-token prediction for faster generation.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"},{"name":"NVFP4 pretraining and weights","detail":"Pre-trained with an NVFP4 recipe; weights published in both BF16 and NVFP4 under the permissive OpenMDW-1.1 license.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"},{"name":"Reasoning on / off / medium","detail":"enable_thinking toggle in the chat template plus a medium-effort mode to cut reasoning tokens.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16"}],"entry":null,"notes":"Knowledge cutoff = pre-training data (Sep 2025); post-training data to May 2026. NVFP4 repo nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4.","verified":"2026-09-29","body_md":"NVIDIA's largest open reasoning model for agentic workflows and long-context analysis.\n\n```bash\ncurl https://integrate.api.nvidia.com/v1/chat/completions -H \"Authorization: Bearer $NVIDIA_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"nvidia/nemotron-3-ultra-550b-a55b\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 , https://integrate.api.nvidia.com/v1/models","page_url":"https://postcutoff.com/m/nemotron-3-ultra/","events_after":834,"major_after":207,"historic_after":41,"missing_at_launch":146,"events_since_release":null,"you_url":"https://postcutoff.com/you/nemotron-3-ultra/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-09.md"}},{"id":"ideogram-4","name":"Ideogram 4.0","org":"Ideogram","family":"Ideogram","released":"2026-06-03","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":true,"model_license":"Ideogram Non-Commercial Model Agreement (quantized open weights); commercial license separate","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Ideogram API","model_id":"ideogram-v4","endpoint":"https://api.ideogram.ai/v1/ideogram-v4/generate","docs":"https://developer.ideogram.ai/api-reference"},{"provider":"Hugging Face","url":"https://huggingface.co/ideogram-ai/ideogram-4-fp8"},{"provider":"Hugging Face (NF4)","url":"https://huggingface.co/ideogram-ai/ideogram-4-nf4"},{"provider":"Web app","url":"https://ideogram.ai"}],"capabilities":[{"name":"Structured JSON prompting with layout control","detail":"Native JSON prompt format with explicit bounding-box layout and color-palette controls.","first":false,"discovered":"launch","source":"https://ideogram.ai/blog/ideogram-4.0/"},{"name":"Best-in-class multilingual text rendering","detail":"Strong in-image typography across languages; native 2K resolution.","first":false,"discovered":"launch","source":"https://ideogram.ai/blog/ideogram-4.0/"},{"name":"First Ideogram open-weight model","detail":"9.3B DiT trained from scratch, Qwen3-VL-8B text encoder; quantized weights on HF for research.","first":false,"discovered":"launch","source":"https://huggingface.co/ideogram-ai/ideogram-4-fp8"}],"entry":null,"notes":"Some third-party sites claim Apache-2.0 - HF card says license: other (non-commercial). API: Api-Key header, multipart with text_prompt or json_prompt; rendering_speed=FLASH currently returns 400. Also /v1/ideogram-v3/generate (previous gen). Per-image API pricing not verified on official page.","verified":"2026-09-29","body_md":"Design-focused image model (posters, logos, layouts with text).\n\n```bash\ncurl -X POST https://api.ideogram.ai/v1/ideogram-v4/generate -H \"Api-Key: $IDEOGRAM_API_KEY\" \\\n  -F text_prompt=\"Minimal concert poster, title 'NIGHT SHIFT' in bold serif\"\n```\n\nSources: https://developer.ideogram.ai/api-reference , https://ideogram.ai/blog/ideogram-4.0/ , https://huggingface.co/collections/ideogram-ai/ideogram-4","page_url":"https://postcutoff.com/m/ideogram-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":688,"you_url":null,"briefings":null},{"id":"mai-voice-2","name":"MAI-Voice-2 / MAI-Voice-2-Flash","org":"Microsoft","family":"MAI-Voice","released":"2026-06-02","status":"preview","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_million_characters":22,"per_million_characters_flash":15,"unit":"USD per 1M characters (MAI-Voice-2 / MAI-Voice-2-Flash, 'starting at')","source":"https://microsoft.ai/models/mai-voice-2/"},"price_line":"$22 per 1M characters","access":[{"provider":"Azure Speech in Microsoft Foundry (SSML voice name)","model_id":"en-US-Harper:MAI-Voice-2","endpoint":"https://{region}.tts.speech.microsoft.com/cognitiveservices/v1","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices"},{"provider":"Azure Speech in Microsoft Foundry (SSML voice name)","model_id":"en-US-Harper:MAI-Voice-2-Flash","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices"},{"provider":"Azure Voice Live (TTS output)","model_id":"MAI-Voice-2-Flash","docs":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/voice-live-how-to"},{"provider":"OpenRouter","model_id":"microsoft/mai-voice-2","endpoint":"https://openrouter.ai/api/v1/audio/speech"},{"provider":"OpenRouter","model_id":"microsoft/mai-voice-2-flash"},{"provider":"Web app (MAI Playground)","url":"https://playground.microsoft.ai/"}],"capabilities":[{"name":"Gated instant voice cloning","detail":"Matches a consented reference voice from a 5-60 s clip without training; only approved (Limited Access) licensed voices can be synthesized.","first":false,"discovered":"launch","source":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices"},{"name":"SSML emotion/style control","detail":"mstts:express-as styles (angry, fearful, joyful, whispering, shouting, etc.) with styledegree, across 15 languages / 18 locales.","first":false,"discovered":"launch","source":"https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices"},{"name":"Low-latency Flash tier","detail":"MAI-Voice-2-Flash (public preview from 2026-07-23) targets voice agents/IVR; Microsoft quotes ~225 ms latency vs ~1 s for MAI-Voice-2 (for a 45 s clip).","first":false,"discovered":"later","source":"https://microsoft.ai/models/mai-voice-2/"}],"entry":"2026-06-02-microsoft-mai-models-build-2026","notes":"Launched at Build 2026-06-02 (MAI-Voice-2); Flash followed 2026-07-23 (date per secondary sources). Both public preview in Azure Speech. Languages include en-US/AU, de, fr, es-ES/MX, pt-BR/PT, it, ko, zh-CN, tr, ru, th, nl, ro, hu, hi. Also used in Copilot (Audio Expressions). Predecessor MAI-Voice-1 no longer listed on the MAI-Voice docs page. Also on Fireworks and Baseten (ids not verified).","verified":"2026-09-29","body_md":"Microsoft's in-house expressive TTS, used through standard Azure Speech SSML.\n\n```bash\ncurl -X POST \"https://$REGION.tts.speech.microsoft.com/cognitiveservices/v1\" \\\n  -H \"Ocp-Apim-Subscription-Key: $SPEECH_KEY\" -H \"Content-Type: application/ssml+xml\" \\\n  -H \"X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3\" \\\n  --data '<speak version=\"1.0\" xml:lang=\"en-US\"><voice name=\"en-US-Harper:MAI-Voice-2\">Hello from MAI Voice.</voice></speak>' -o out.mp3\n```\n\nSources: https://learn.microsoft.com/en-us/azure/ai-services/speech-service/mai-voices · https://microsoft.ai/models/mai-voice-2/ · https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/","page_url":"https://postcutoff.com/m/mai-voice-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":688,"you_url":null,"briefings":null},{"id":"cosmos-3","name":"Cosmos 3 (Nano / Super)","org":"NVIDIA","family":"Cosmos","released":"2026-06-01","status":"current","type":"world-model","modality_in":["text","image","video","audio","action"],"modality_out":["text","image","video","audio","action"],"open_weights":true,"model_license":"OpenMDW-1.1 (commercial and non-commercial use)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/Cosmos3-Nano","url":"https://huggingface.co/nvidia/Cosmos3-Nano"},{"provider":"Hugging Face","model_id":"nvidia/Cosmos3-Super","url":"https://huggingface.co/nvidia/Cosmos3-Super"},{"provider":"GitHub","url":"https://github.com/nvidia-cosmos","docs":"https://research.nvidia.com/labs/cosmos-lab/cosmos3/"}],"capabilities":[{"name":"Unified omni world model (generation + reasoning + action)","detail":"One Mixture-of-Transformers model (autoregressive + diffusion) replaces separate Cosmos Predict, Transfer, Reason and Policy models: world generation, physical reasoning, forward/inverse dynamics and action/policy generation.","first":true,"discovered":"launch","source":"https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world"},{"name":"Open omnimodal I/O","detail":"Inputs text, images, short video, audio and action trajectories (16-400 frames); outputs text, images, video (5-400 frames), 48 kHz stereo audio and actions (JSON).","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/Cosmos3-Nano"},{"name":"Leaderboard results","detail":"NVIDIA cites best open text-to-image and image-to-video models on Artificial Analysis and best policy model on RoboArena.","first":false,"discovered":"later","source":"https://www.nvidia.com/en-us/ai/cosmos/"}],"entry":"2026-06-01-nvidia-cosmos-3-open-release","notes":"Announced at GTC 2026-03-16 ('the first world foundation model unifying synthetic world generation, vision reasoning and action simulation' - NVIDIA claim); weights published 2026-05-31/06-01 (HF blog 'The First Open Omni-model for Physical AI Reasoning and Action'). Sizes: Nano 16B, Super 64B. Linux + Ampere/Hopper/Blackwell GPUs, BF16. Technical report dated 2026-06-22.","verified":"2026-09-29","body_md":"Sources: [HF blog](https://huggingface.co/blog/nvidia/cosmos-3-for-physical-ai), [Cosmos3-Nano](https://huggingface.co/nvidia/Cosmos3-Nano), [Cosmos3-Super](https://huggingface.co/nvidia/Cosmos3-Super), [technical report](https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf).","page_url":"https://postcutoff.com/m/cosmos-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":691,"you_url":null,"briefings":null},{"id":"minimax-m3","name":"MiniMax-M3","org":"MiniMax","family":"MiniMax-M","released":"2026-06-01","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"minimax-community","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":1.2,"cache_read":0.06,"unit":"per 1M tokens (USD), standard tier, input <=512K (after permanent 50% discount); >512K input: 0.60/2.40/0.12. Priority tier 1.5x","source":"https://platform.minimax.io/docs/guides/pricing-paygo"},"price_line":"$0.30 in, $1.20 out per 1M tokens","access":[{"provider":"MiniMax API (Anthropic format)","model_id":"MiniMax-M3","endpoint":"https://api.minimax.io/anthropic","docs":"https://platform.minimax.io/docs/api-reference/text-anthropic-api"},{"provider":"MiniMax API (OpenAI format)","model_id":"MiniMax-M3","endpoint":"https://api.minimax.io/v1","docs":"https://platform.minimax.io/docs/api-reference/text-openai-api"},{"provider":"OpenRouter","model_id":"minimax/minimax-m3","url":"https://openrouter.ai/minimax/minimax-m3"},{"provider":"Hugging Face","url":"https://huggingface.co/MiniMaxAI/MiniMax-M3"}],"capabilities":[{"name":"MiniMax Sparse Attention (MSA)","detail":"New sparse attention for million-token contexts: 9x prefill and 15x decode speed-up vs M2 at 1M context, ~1/20 per-token compute.","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-M3"},{"name":"Native multimodality from step one","detail":"Mixed text/image/video training from the start of pre-training (~428B total / ~23B active).","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-M3"},{"name":"Three reasoning modes","detail":"thinking parameter selects among three reasoning modes; interleaved thinking with tool use.","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-M3"}],"entry":null,"notes":"MiniMax flagship LLM. OpenAI-format responses include <think> content that must be preserved across turns. MiniMax-M3.1-Flash-Preview (1M, tunable thinking) exists but only via Token Plan/MiniMax Code. Max output and knowledge cutoff not verified.","verified":"2026-09-29","body_md":"Low-cost 1M-context multimodal coding/agent model; Anthropic-SDK-first API, open weights.\n\n```bash\ncurl https://api.minimax.io/anthropic/v1/messages \\\n -H \"x-api-key: $MINIMAX_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"MiniMax-M3\",\"max_tokens\":4096,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://huggingface.co/MiniMaxAI/MiniMax-M3","page_url":"https://postcutoff.com/m/minimax-m3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":691,"you_url":"https://postcutoff.com/you/minimax-m3/","briefings":null},{"id":"kimi-k2-7-code","name":"Kimi K2.7 Code","org":"Moonshot AI","family":"Kimi K2","released":"2026-06","status":"current","type":"code","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"modified-mit","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.95,"output":4,"cache_read":0.19,"unit":"per 1M tokens (USD); kimi-k2.7-code-highspeed: 1.90 in / 8.00 out / 0.38 cache hit","source":"https://platform.kimi.ai/docs/pricing/chat"},"price_line":"$0.95 in, $4 out per 1M tokens","access":[{"provider":"Kimi API (Moonshot)","model_id":"kimi-k2.7-code","endpoint":"https://api.moonshot.ai/v1","docs":"https://platform.kimi.ai/docs/guide/kimi-k2-7-code-quickstart"},{"provider":"Kimi API (Moonshot) high-speed","model_id":"kimi-k2.7-code-highspeed","endpoint":"https://api.moonshot.ai/v1","docs":"https://platform.kimi.ai/docs/models"},{"provider":"OpenRouter","model_id":"moonshotai/kimi-k2.7-code","url":"https://openrouter.ai/moonshotai/kimi-k2.7-code"},{"provider":"Hugging Face","url":"https://huggingface.co/moonshotai/Kimi-K2.7-Code"},{"provider":"Web app","url":"https://www.kimi.com"}],"capabilities":[{"name":"Coding-specialized K2.6 derivative","detail":"Built on Kimi K2.6 (1T total / 32B active, MLA, 400M vision encoder) and tuned for long-horizon real-world coding.","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2.7-Code"},{"name":"~30% fewer thinking tokens than K2.6","detail":"Higher task success with about 30% lower thinking-token usage vs K2.6.","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2.7-Code"},{"name":"High-speed tier","detail":"kimi-k2.7-code-highspeed outputs ~180 tok/s (up to ~260 tok/s on short context).","first":false,"discovered":"launch","source":"https://platform.kimi.ai/docs/models"}],"entry":null,"notes":"Dedicated coding model; pairs with Kimi Code CLI. Release day not verified (HF 2026-06-11, OpenRouter 2026-06-12). Max output not verified.","verified":"2026-09-29","body_md":"Cheaper-than-K3 coding agent model for Kimi Code CLI, Claude Code, Codex, OpenCode.\n\n```bash\ncurl https://api.moonshot.ai/v1/chat/completions \\\n -H \"Authorization: Bearer $MOONSHOT_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"kimi-k2.7-code\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a Python quicksort\"}]}'\n```\n\nSources: https://platform.kimi.ai/docs/models · https://platform.kimi.ai/docs/pricing/chat · https://huggingface.co/moonshotai/Kimi-K2.7-Code","page_url":"https://postcutoff.com/m/kimi-k2-7-code/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":665,"you_url":null,"briefings":null},{"id":"vui-luna-tts","name":"VUI Labs Luna-TTS (and Luna-TTS Realtime)","org":"VUI Labs","family":"Luna-TTS","released":"2026-06","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":15,"unit":"USD per 1M characters (Luna-TTS and Luna-TTS Character); voice cloning $3 per clone","source":"https://www.vuilabs.ai/"},"price_line":"$15 per 1M characters","access":[{"provider":"VUI Labs API","url":"https://www.vuilabs.ai/"},{"provider":"arXiv (technical report)","url":"https://arxiv.org/abs/2608.11593"}],"capabilities":[{"name":"Diffusion-language-model TTS (non-autoregressive)","detail":"Generates the whole RVQ token grid in a fixed number of parallel refinement steps; the Realtime variant is blockwise-autoregressive over 1.28 s blocks (RTF 0.0240, 41.6 ms first-block latency locally). 0.6B backbone, ~1M hours of zh/en/ja/ko speech.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2608.11593"},{"name":"Chinese startup at the top of TTS arenas","detail":"Pandaily (Aug 2026) reported #1 on Hugging Face TTS Arena and #3 on Artificial Analysis Speech Arena; on 2026-09-29 AA shows it #8 (Elo 1230).","first":false,"discovered":"later","source":"https://pandaily.com/vui-labs-luna-tts-number-one-tts-arena-qian-yanmin-voice-agent-aug2026"}],"entry":null,"notes":"Chinese voice-AI startup (Pandaily). Release month June 2026 per the Artificial Analysis leaderboard; technical report 2026-08-12 (Feng Yin et al., 22 authors). Supports zero-shot cloning, speech editing, emotion control, non-verbal vocalisations. We found no statement about open weights. Pandaily headline calls it China's 'Thinking Machines' and names Qian Yanmin (role not verified). Not the same as fluxions-ai 'Vui' (open Apache-2.0 small TTS).","verified":"2026-09-29","body_md":"Sources: https://arxiv.org/abs/2608.11593 , https://www.vuilabs.ai/ , https://artificialanalysis.ai/text-to-speech/model-families/luna-tts","page_url":"https://postcutoff.com/m/vui-luna-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":665,"you_url":null,"briefings":null},{"id":"grok-imagine-video-1-5","name":"Grok Imagine Video 1.5","org":"xAI","family":"Grok Imagine","released":"2026-05-30","status":"current","type":"video-gen","modality_in":["text","image"],"modality_out":["video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.08,"unit":"per second of generated video","source":"https://docs.x.ai/docs/models"},"price_line":"$0.08 per second","access":[{"provider":"xAI API","model_id":"grok-imagine-video-1.5","endpoint":"https://api.x.ai/v1/videos/generations","docs":"https://docs.x.ai/docs/guides/video-generation"},{"provider":"Web app","url":"https://grok.com"}],"capabilities":[{"name":"Image-to-video up to 15 s","detail":"Animates a source still (URL/base64) or prompt into clips up to 15 seconds; async job polled via GET /v1/videos/{request_id}.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/guides/video-generation"},{"name":"Per-second pricing, text or image input","detail":"Text- or image-to-video at $0.08 per generated second (legacy grok-imagine-video $0.05/s).","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models"}],"entry":null,"notes":"Snapshot alias grok-imagine-video-1.5-2026-05-30. Legacy grok-imagine-video still available at $0.05/s. Resolution/audio details not verified.","verified":"2026-09-29","body_md":"Short video generation from text or an image.\n\n```bash\ncurl -X POST https://api.x.ai/v1/videos/generations -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-imagine-video-1.5\",\"prompt\":\"A paper boat drifting down a rainy street\",\"duration\":8}'\n# then poll: curl https://api.x.ai/v1/videos/$REQUEST_ID -H \"Authorization: Bearer $XAI_API_KEY\"\n```\n\nSources: https://docs.x.ai/docs/models/grok-imagine-video-1.5 , https://docs.x.ai/docs/guides/video-generation","page_url":"https://postcutoff.com/m/grok-imagine-video-1-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":695,"you_url":null,"briefings":null},{"id":"claude-opus-4-8","name":"Claude Opus 4.8","org":"Anthropic","family":"Claude 4","released":"2026-05-28","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-01","pricing":{"input":5,"output":25,"cache_read":0.5,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-4-8","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-4-8/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-opus-4-8","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-4-8","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-4-8","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-4-8","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-4.8","url":"https://openrouter.ai/anthropic/claude-opus-4.8"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Code honesty","detail":"About 4x less likely than Opus 4.7 to let flaws in its own code pass without comment.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-8"},{"name":"Browser agents","detail":"Scored 84% on Online-Mind2Web, ahead of Opus 4.7 and GPT-5.5.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-8"},{"name":"Legal agent benchmark","detail":"Anthropic says it was the first model to exceed 10% on the Legal Agent Benchmark all-pass standard.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-8"},{"name":"Cheaper fast mode","detail":"Fast mode runs up to 2.5x faster, at a lower premium than earlier fast mode.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-8"}],"entry":null,"notes":"Last Opus 4.x. Adaptive thinking only (omit thinking = no thinking); sampling params and budget_tokens removed. Fast mode $10/$50 (Claude API only). Retirement not sooner than 2027-05-28.","verified":"2026-09-29","body_md":"Legacy; common fallback target for refusals on Claude 5 models.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-4-8\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-4-8/overview\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-4-8/","events_after":793,"major_after":197,"historic_after":40,"missing_at_launch":94,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-4-8/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-01.md"}},{"id":"elevenlabs-dubbing-v2","name":"Eleven Dubbing v2","org":"ElevenLabs","family":"Eleven Dubbing","released":"2026-05-28","status":"preview","type":"audio/speech","modality_in":["audio","video"],"modality_out":["audio","video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":2.2,"unit":"per minute (Dubbing v2 API); Dubbing v1 $0.33/min watermarked, $0.50/min unwatermarked","source":"https://elevenlabs.io/pricing/api"},"price_line":"$2.20 per minute","access":[{"provider":"ElevenLabs API","endpoint":"https://api.elevenlabs.io/v1/dubbing","docs":"https://elevenlabs.io/docs/overview/capabilities/dubbing"},{"provider":"Web app (ElevenCreative / ElevenProductions)","url":"https://elevenlabs.io/dubbing"}],"capabilities":[{"name":"Direct speech-to-speech dubbing","detail":"Conditions directly on the original performance instead of an ASR -> translate -> TTS pipeline, so intonation and emotion carry across 90+ languages; ElevenLabs: 'For the first time, the emotion and performance of the original speaker carries across every language' (company claim, not independently verified as a first).","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/introducing-dubbing-v2"},{"name":"Project-based dubbing API","detail":"API (2026-08-06/10) with editable JSON transcripts/translations, regional variants (e.g. es-MX), sync-aware translation; 3 GB per file via API.","first":false,"discovered":"later","source":"https://elevenlabs.io/blog/dubbing-api"}],"entry":"2026-05-28-elevenlabs-dubbing-v2","notes":"Launched in UI 2026-05-28; API announced 2026-08-06 (blog) / changelog 2026-08-10. Docs label it 'Dubbing v2 Alpha' (default for Automatic Dubbing), hence status preview. No explicit model_id string found in API docs.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/blog/introducing-dubbing-v2 , https://elevenlabs.io/blog/dubbing-api , https://elevenlabs.io/docs/overview/capabilities/dubbing , https://elevenlabs.io/pricing/api , https://x.com/ElevenLabsDevs/status/2085380402508619880","page_url":"https://postcutoff.com/m/elevenlabs-dubbing-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":695,"you_url":null,"briefings":null},{"id":"qwen3-7-plus","name":"Qwen3.7-Plus","org":"Alibaba (Qwen)","family":"Qwen3.7","released":"2026-05-26","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.4,"output":1.6,"unit":"per 1M tokens (USD), Singapore/International, input up to 256K (list price; limited-time 20% off). 256K-1M input: 1.2 / 4.8","source":"https://www.alibabacloud.com/help/en/model-studio/model-pricing"},"price_line":"$0.40 in, $1.60 out per 1M tokens","access":[{"provider":"Alibaba Cloud Model Studio (DashScope, Singapore/Intl)","model_id":"qwen3.7-plus","endpoint":"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus"},{"provider":"OpenRouter","model_id":"qwen/qwen3.7-plus","url":"https://openrouter.ai/qwen/qwen3.7-plus"},{"provider":"Web app","url":"https://chat.qwen.ai"}],"capabilities":[{"name":"Multimodal hybrid GUI agent","detail":"Perceives real-world scenes, reads screens and operates GUIs, generates code from visual references and navigates mobile apps end to end.","first":false,"discovered":"launch","source":"https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus"},{"name":"Recommended balanced coding model","detail":"Alibaba's recommended model for coding tools: full tool calling, built-in tools and 1M context at mid-tier price.","first":false,"discovered":"later","source":"https://www.alibabacloud.com/help/en/model-studio/text-generation-model"}],"entry":null,"notes":"Alias of snapshot qwen3.7-plus-2026-05-26 (release date taken from the snapshot name). Thinking and non-thinking modes. Max output not verified.","verified":"2026-09-29","body_md":"Balanced price/performance Qwen for chatbots, document processing and coding agents; strong GUI/vision-agent skills.\n\n```bash\ncurl \"https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1/chat/completions\" \\\n -H \"Authorization: Bearer $DASHSCOPE_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"qwen3.7-plus\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://www.alibabacloud.com/help/en/model-studio/qwen3-7-plus · https://www.alibabacloud.com/help/en/model-studio/text-generation-model · https://www.alibabacloud.com/help/en/model-studio/model-pricing","page_url":"https://postcutoff.com/m/qwen3-7-plus/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":700,"you_url":"https://postcutoff.com/you/qwen3-7-plus/","briefings":null},{"id":"runway-aleph-2","name":"Runway Aleph 2.0","org":"Runway","family":"Aleph","released":"2026-05-21","status":"current","type":"video-gen","modality_in":["video","text","image"],"modality_out":["video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.28,"unit":"per second of video (28 credits/s at $0.01/credit, 56-credit minimum per generation)","source":"https://docs.dev.runwayml.com/guides/pricing/"},"price_line":"$0.28 per second","access":[{"provider":"Runway API","model_id":"aleph2","endpoint":"https://api.dev.runwayml.com/v1/video_to_video","docs":"https://docs.dev.runwayml.com/guides/models/"},{"provider":"Web app","url":"https://app.runwayml.com"}],"capabilities":[{"name":"In-context video editing of real footage","detail":"Edits existing clips (up to 30 s of 1080p): change angles, lighting, objects, wardrobe, background while preserving untouched motion and scene structure.","first":false,"discovered":"launch","source":"https://runway.com/news/introducing-aleph-2-and-edit-studio"},{"name":"Edit one frame, propagate to the clip","detail":"Image-level keyframe control (up to 5 keyframes in the API) and multi-shot edits applied across scene cuts.","first":false,"discovered":"launch","source":"https://docs.dev.runwayml.com/api-details/api_changelog/"}],"entry":null,"notes":"Launched with Edit Studio 2026-05-21; API since 2026-06-02 (2-30 s input videos). Supersedes gen4_aleph (removed from API 2026-07-30).","verified":"2026-09-29","body_md":"Video-to-video editing model: prompt-driven edits on real footage.\n\n```bash\ncurl -X POST https://api.dev.runwayml.com/v1/video_to_video -H \"Authorization: Bearer $RUNWAYML_API_SECRET\" \\\n  -H \"X-Runway-Version: 2024-11-06\" -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"aleph2\",\"videoUri\":\"https://.../clip.mp4\",\"promptText\":\"make it night with neon rain\"}'\n```\n\nSources: https://runway.com/news/introducing-aleph-2-and-edit-studio , https://docs.dev.runwayml.com/guides/models/ , https://docs.dev.runwayml.com/guides/pricing/","page_url":"https://postcutoff.com/m/runway-aleph-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":701,"you_url":null,"briefings":null},{"id":"command-a-plus","name":"Command A+","org":"Cohere","family":"Command A","released":"2026-05-20","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":128000,"max_output":64000,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":1.5,"unit":"per 1M tokens (USD) on OpenRouter; Cohere first-party price not verified","source":"https://openrouter.ai/cohere/command-a-plus"},"price_line":"$0.30 in, $1.50 out per 1M tokens","access":[{"provider":"Cohere API","model_id":"command-a-plus-05-2026","endpoint":"https://api.cohere.com/v2/chat","docs":"https://docs.cohere.com/docs/command-a-plus"},{"provider":"OpenRouter","model_id":"cohere/command-a-plus","url":"https://openrouter.ai/cohere/command-a-plus"},{"provider":"Hugging Face","url":"https://huggingface.co/CohereLabs/command-a-plus-05-2026-w4a4"}],"capabilities":[{"name":"Cohere's first MoE model","detail":"218B total / 25B active mixture-of-experts combining vision, agentic and reasoning capabilities in one model.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/models"},{"name":"Apache 2.0 enterprise model on 1 B200","detail":"Open weights under Apache 2.0 (earlier Command A was CC-BY-NC); W4A4 build runs on 1x B200 or 2x H100.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/command-a-plus"},{"name":"48 languages","detail":"Supports 48 languages including all official EU languages, with configurable reasoning.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/command-a-plus"}],"entry":null,"notes":"Also HF CohereLabs/command-a-plus-05-2026-bf16 and -fp8. OpenRouter lists 192K context vs 128K in Cohere docs. Cohere pricing page did not list per-token price.","verified":"2026-09-29","body_md":"Cohere's current flagship for enterprise agents, RAG, multilingual and vision.\n\n```bash\ncurl https://api.cohere.com/v2/chat -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"command-a-plus-05-2026\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.cohere.com/docs/models , https://docs.cohere.com/docs/command-a-plus","page_url":"https://postcutoff.com/m/command-a-plus/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":702,"you_url":"https://postcutoff.com/you/command-a-plus/","briefings":null},{"id":"stable-audio-3","name":"Stable Audio 3.0","org":"Stability AI","family":"Stable Audio","released":"2026-05-20","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"Stable Audio Community License (Small SFX / Small / Medium); Large is API/enterprise only","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Stability AI API (Large)","docs":"https://platform.stability.ai"},{"provider":"Hugging Face (Medium)","url":"https://huggingface.co/stabilityai/stable-audio-3-medium"},{"provider":"Hugging Face (Small music)","url":"https://huggingface.co/stabilityai/stable-audio-3-small-music"},{"provider":"Hugging Face (Small SFX)","url":"https://huggingface.co/stabilityai/stable-audio-3-small-sfx"},{"provider":"Web app","url":"https://stableaudio.com"}],"capabilities":[{"name":"Tracks over 6 minutes","detail":"Medium generates music up to 6:20; Large aimed at high-volume, low-latency platform use.","first":false,"discovered":"launch","source":"https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models"},{"name":"Fully licensed training data","detail":"Model family trained on fully licensed data; users own outputs under the Community License.","first":false,"discovered":"launch","source":"https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models"},{"name":"On-device small models","detail":"Small (459M) music and Small SFX models designed to run on phones and consumer laptops.","first":false,"discovered":"launch","source":"https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models"}],"entry":null,"notes":"Family of 4: Small SFX, Small, Medium (open weights, HF) and Large (API via Stability and fal.ai, or enterprise self-hosting). Exact API model id/endpoint for Large not verified (Stability pricing/docs pages are JS-rendered). Price per The Rundown tool review (says it checked the official pricing page 2026-08-31, secondary): 26 API credits = $0.26 per successful Large generation (1 credit = $0.01).","verified":"2026-09-29","body_md":"Stability's licensed-data music & sound-effects family; open Medium/Small checkpoints for local use.\n\nSources: https://stability.ai/news-updates/meet-stable-audio-3-the-model-family-built-for-artistic-experimentation-with-open-weight-models , https://huggingface.co/collections/stabilityai/stable-audio-3 , https://github.com/Stability-AI/stable-audio-3\n\n## Changelog\n- 2026-09-29: added secondary-source Large API price (26 credits/$0.26 per generation, https://www.therundown.ai/tools/stable-audio-3-0); model id still unverified","page_url":"https://postcutoff.com/m/stable-audio-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":702,"you_url":null,"briefings":null},{"id":"gemini-3-5-flash","name":"Gemini 3.5 Flash","org":"Google DeepMind","family":"Gemini 3.5","released":"2026-05-19","status":"legacy","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":1.5,"output":9,"unit":"per 1M tokens (Standard tier)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1.50 in, $9 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.5-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"gemini-3.5-flash","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-5-flash"},{"provider":"OpenRouter","model_id":"google/gemini-3.5-flash"}],"capabilities":[{"name":"Flash beats previous Pro on agents","detail":"Outperformed Gemini 3.1 Pro on Terminal-Bench 2.1 (76.2%), GDPval-AA (1656 Elo) and MCP Atlas (83.6%); 84.2% CharXiv Reasoning.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/"},{"name":"High output speed","detail":"Google claims ~4x the output tokens/second of other frontier models at launch.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/"}],"entry":null,"notes":"Launched at Google I/O 2026 as first Gemini 3.5 model. Model page lists gemini-3-flash-preview (Dec 2025) as its preview predecessor id; that preview is still served. Now more expensive than 3.6-3.8 Flash; use gemini-3.8-flash.","verified":"2026-09-29","body_md":"First model of the Gemini 3.5 generation (I/O, May 2026). Kept for compatibility; newer 3.6-3.8 Flash models are cheaper and stronger.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [I/O blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/).","page_url":"https://postcutoff.com/m/gemini-3-5-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":705,"you_url":null,"briefings":null},{"id":"recraft-v4-1","name":"Recraft V4.1","org":"Recraft","family":"Recraft V4","released":"2026-05-14","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.035,"unit":"per image (recraftv4_1 raster; Pro 0.21, Vector 0.08, V4.1 Flash 0.007)","source":"https://www.recraft.ai/docs/api-reference/pricing"},"price_line":"$0.035 per image","access":[{"provider":"Recraft API","model_id":"recraftv4_1","endpoint":"https://external.api.recraft.ai/v1/images/generations","docs":"https://www.recraft.ai/docs/api-reference/getting-started"},{"provider":"OpenRouter","model_id":"recraft/recraft-v4.1"},{"provider":"Web app","url":"https://www.recraft.ai"}],"capabilities":[{"name":"Native vector (SVG) generation","detail":"Dedicated Vector variants (recraftv4_1_vector, _pro_vector) output editable vector logos, typography and illustrations.","first":false,"discovered":"launch","source":"https://www.recraft.ai/blog/recraft-v4-1-more-beautiful-by-nature"},{"name":"Utility variant for mockups","detail":"V4.1 Utility gives flat lighting, front-facing product/mockup outputs alongside the expressive main model.","first":false,"discovered":"launch","source":"https://www.recraft.ai/blog/recraft-v4-1-more-beautiful-by-nature"},{"name":"V4.1 Flash","detail":"Sept 2026 fast variant (~1.3 s end-to-end) at $0.007/image.","first":false,"discovered":"later","source":"https://www.recraft.ai/docs/api-reference/getting-started"}],"entry":null,"notes":"Model ids: recraftv4_1, recraftv4_1_pro, recraftv4_1_vector, recraftv4_1_pro_vector, recraftv4_1_utility(_pro)(_vector), recraftv4_1_flash; earlier recraftv4 ($0.04), recraftv4_styles, recraftv3. OpenAI-SDK compatible.","verified":"2026-09-29","body_md":"Design-oriented image model family (raster + vector). OpenAI-compatible API.\n\n```python\nfrom openai import OpenAI\nc = OpenAI(base_url=\"https://external.api.recraft.ai/v1\", api_key=RECRAFT_API_TOKEN)\nc.images.generate(model=\"recraftv4_1\", prompt=\"flat vector logo of a fox, orange and navy\")\n```\n\nSources: https://www.recraft.ai/docs/api-reference/getting-started , https://www.recraft.ai/docs/api-reference/pricing , https://www.recraft.ai/blog/recraft-v4-1-more-beautiful-by-nature , https://openrouter.ai/recraft/recraft-v4.1","page_url":"https://postcutoff.com/m/recraft-v4-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":710,"you_url":null,"briefings":null},{"id":"gemini-3-1-flash-lite","name":"Gemini 3.1 Flash-Lite","org":"Google DeepMind","family":"Gemini 3","released":"2026-05-07","status":"deprecated","type":"llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":0.25,"output":1.5,"unit":"per 1M tokens (Standard; text/image/video input; audio input $0.50)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.25 in, $1.50 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.1-flash-lite","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"},{"provider":"OpenRouter","model_id":"google/gemini-3.1-flash-lite"}],"capabilities":[{"name":"Low-cost frontier-class Lite","detail":"Described as frontier-class performance at reduced cost; cheapest per-token 3.x text model.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models"},{"name":"Full tool stack on a Lite model","detail":"1M-token multimodal input (text, image, video, audio, PDF) with 65K output.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite"}],"entry":null,"notes":"Stable GA 2026-05-07; scheduled shutdown 2027-05-07, replacement gemini-3.5-flash-lite. Preview id gemini-3.1-flash-lite-preview (early 2026) still listed by the live API / OpenRouter though docs list it as shut down.","verified":"2026-09-29","body_md":"Budget Gemini 3.1 model; migrate to `gemini-3.5-flash-lite` before May 2027.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/gemini-3-1-flash-lite/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":716,"you_url":null,"briefings":null},{"id":"gpt-realtime-2","name":"GPT-Realtime-2","org":"OpenAI","family":"GPT Realtime","released":"2026-05-07","status":"current","type":"audio/speech","modality_in":["text","audio","image"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":32000,"knowledge_cutoff":"2024-09","pricing":{"text_input":4,"text_output":24,"audio_input":32,"audio_output":64,"cached_input":0.4,"image_input":5,"unit":"per 1M tokens (USD); cached audio/text input $0.40, cached image $0.50","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$4 text in, $24 text out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-realtime-2","endpoint":"wss://api.openai.com/v1/realtime","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-2"}],"capabilities":[{"name":"Reasoning speech-to-speech model","detail":"First OpenAI realtime voice model with configurable reasoning effort (press: 'GPT-5-class' reasoning); higher effort adds latency and tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-2"},{"name":"128K-token realtime context","detail":"Context grew from 32K (gpt-realtime-1.5) to 128K tokens, with 32K max output, for long voice-agent sessions.","first":false,"discovered":"launch","source":"https://www.ghacks.net/2026/05/11/openai-releases-three-new-realtime-voice-models-for-the-api-with-gpt-5-class-reasoning/"}],"entry":"2026-05-07-openai-gpt-realtime-2-translate-whisper","notes":"Launched 2026-05-07 with gpt-realtime-translate and gpt-realtime-whisper (changelog). Superseded two months later by gpt-realtime-2.1 (2026-07-06) at identical prices, but still listed and not deprecated. Realtime endpoint only; function calling and prompt caching. Official launch post (openai.com) returned 403 to our fetcher, so benchmark claims were not read directly; secondary sources quote OpenAI: +15.2% Big Bench Audio vs gpt-realtime-1.5 (high effort), +13.8% Audio MultiChallenge instruction following (xhigh); one blog reports 96.6% absolute Big Bench Audio at xhigh (unconfirmed).","verified":"2026-09-29","body_md":"OpenAI's first reasoning realtime voice model. For new builds prefer [gpt-realtime-2.1](gpt-realtime-2-1.md).\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-realtime-2\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/","page_url":"https://postcutoff.com/m/gpt-realtime-2/","events_after":897,"major_after":238,"historic_after":51,"missing_at_launch":179,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-09.md"}},{"id":"gpt-realtime-translate","name":"GPT-Realtime-Translate","org":"OpenAI","family":"GPT Realtime","released":"2026-05-07","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":16000,"max_output":2000,"knowledge_cutoff":null,"pricing":{"per_minute":0.034,"unit":"per minute of audio (USD), billed by duration not tokens","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.034 per minute","access":[{"provider":"OpenAI API","model_id":"gpt-realtime-translate","endpoint":"https://api.openai.com/v1/realtime/translations","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-translate"}],"capabilities":[{"name":"Streaming speech-to-speech translation","detail":"Simultaneous interpretation from 70+ input languages into 13 output languages, emitting translated audio plus transcript deltas while the speaker is still talking.","first":false,"discovered":"launch","source":"https://www.ghacks.net/2026/05/11/openai-releases-three-new-realtime-voice-models-for-the-api-with-gpt-5-class-reasoning/"},{"name":"Dedicated translation endpoint","detail":"Served only on v1/realtime/translations (not the general Realtime or Chat endpoints).","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-translate"}],"entry":"2026-05-07-openai-gpt-realtime-2-translate-whisper","notes":"Language counts (70+ in / 13 out) come from press coverage of the launch post; the docs page does not list languages. Latency not specified. Google's comparable model is gemini-3.5-live-translate-preview (June 2026).","verified":"2026-09-29","body_md":"Live interpretation for calls, events and apps.\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-realtime-translate\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog","page_url":"https://postcutoff.com/m/gpt-realtime-translate/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":716,"you_url":null,"briefings":null},{"id":"gpt-realtime-whisper","name":"GPT-Realtime-Whisper","org":"OpenAI","family":"GPT Realtime","released":"2026-05-07","status":"legacy","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":16000,"max_output":2000,"knowledge_cutoff":null,"pricing":{"per_minute":0.017,"unit":"per minute of audio (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.017 per minute","access":[{"provider":"OpenAI API","model_id":"gpt-realtime-whisper","endpoint":"wss://api.openai.com/v1/realtime (transcription sessions, v1/realtime/transcription_sessions)","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-whisper"}],"capabilities":[{"name":"Streaming speech-to-text with tunable latency","detail":"Streams transcript deltas from live audio with a latency/accuracy trade-off setting.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-whisper"}],"entry":"2026-05-07-openai-gpt-realtime-2-translate-whisper","notes":"Still listed and not deprecated, but gpt-live-transcribe (2026-07-28, same $0.017/min) adds context and keyword hints and is what OpenAI recommends in its deprecation notices; hence marked legacy here. Language list not given in docs.","verified":"2026-09-29","body_md":"Streaming transcription for live audio; see [gpt-live-transcribe](gpt-live-transcribe.md) for the newer option.\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-realtime-whisper\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog","page_url":"https://postcutoff.com/m/gpt-realtime-whisper/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":716,"you_url":null,"briefings":null},{"id":"molmoact-2","name":"MolmoAct 2 / MolmoAct 2-Think","org":"Ai2 (Allen Institute for AI)","family":"MolmoAct","released":"2026-05-05","status":"current","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":true,"model_license":"Apache-2.0 (code); model weights on HF (license tag not stated on card)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"allenai/MolmoAct2","url":"https://huggingface.co/collections/allenai/molmoact2-models"},{"provider":"GitHub","url":"https://github.com/allenai/molmoact2"},{"provider":"Hugging Face LeRobot","model_id":"allenai/MolmoAct2-LIBERO-LeRobot","url":"https://huggingface.co/allenai/MolmoAct2-LIBERO-LeRobot"}],"capabilities":[{"name":"Open action reasoning model","detail":"Molmo2-ER embodied-reasoning VLM connected to a flow-matching action expert via per-layer KV conditioning; the Think variant adds adaptive depth reasoning (interpretable depth map before acting).","first":false,"discovered":"launch","source":"https://allenai.org/blog/molmoact2"},{"name":"Strong out-of-the-box real-world success","detail":"87.1% average success over 15 real Franka tasks vs 45.2% for π0.5 and 48.4% for MolmoBot (Ai2's evaluation); LIBERO 97.2% (98.1% Think).","first":false,"discovered":"launch","source":"https://allenai.org/blog/molmoact2"},{"name":"Fast inference","detail":"~180 ms per action call (790 ms with adaptive depth reasoning) vs ~6,700 ms for the original MolmoAct (up to 37x faster).","first":false,"discovered":"launch","source":"https://allenai.org/blog/molmoact2"},{"name":"Largest open bimanual dataset","detail":"Released with MolmoAct2-BimanualYAM, 720+ hours of bimanual tabletop demonstrations, which Ai2 calls the largest open bimanual robotics dataset, plus an open FAST action tokenizer.","first":false,"discovered":"launch","source":"https://allenai.org/blog/molmoact2"}],"entry":"2026-05-05-ai2-molmoact-2","notes":"Checkpoints: MolmoAct2 (post-trained multi-embodiment foundation, ~5.4B params per HF safetensors), -Think, -Pretrain, fine-tuned -DROID, -BimanualYAM, -SO100_101, -LIBERO, -Think-LIBERO, FAST-Tokenizer. Main supported robots: SO-100/101, bimanual YAM, Franka (DROID); others need fine-tuning. Paper arXiv 2605.02881.","verified":"2026-09-29","body_md":"Fully open (weights, data, code) VLA from Ai2, the main open alternative to π0.5/GR00T for tabletop manipulation. Start from a fine-tuned checkpoint (e.g. `allenai/MolmoAct2-DROID`) for ready-to-run inference; the base card has no inference code.\n\nSources: [Ai2 blog](https://allenai.org/blog/molmoact2), [arXiv 2605.02881](https://arxiv.org/abs/2605.02881), [HF model card](https://huggingface.co/allenai/MolmoAct2), [GitHub](https://github.com/allenai/molmoact2).","page_url":"https://postcutoff.com/m/molmoact-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":719,"you_url":null,"briefings":null},{"id":"gemini-omni-flash","name":"Gemini Omni Flash (Omni 1.1 Flash)","org":"Google DeepMind","family":"Gemini Omni","released":"2026-05","status":"current","type":"video-gen","modality_in":["text","image","audio","video"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1.5,"output":9,"per_second_video":0.1,"unit":"per 1M tokens for text/image/video/audio input ($1.50) and text output ($9.00); video output $17.50 per 1M tokens (~$0.10 per second at 720p)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1.50 in, $9 out per 1M tokens","access":[{"provider":"Gemini API (Interactions API)","model_id":"gemini-omni-1.1-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/omni-1-1-flash"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Any-input video generation","detail":"Generates video with native audio from any mix of text, image, audio and video input, grounded in Gemini world knowledge.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/"},{"name":"Conversational video editing","detail":"Edit, extend (inputs up to 10 s), interpolate keyframes and upscale videos through multi-turn natural-language conversation.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash"},{"name":"Up to 4K output","detail":"3-10 s clips at 360p/720p/1080p/4K, 24 fps.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash"},{"name":"Avatars + SynthID","detail":"Launched with avatar support (your own digital likeness); all outputs carry SynthID watermarks.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/"}],"entry":null,"notes":"Announced at Google I/O 2026 (preview id gemini-omni-flash-preview); Omni 1.1 Flash GA in the API 2026-08-27. Google's recommended default video model over Veo 3.1. Live API model list reports 131k context for gemini-omni-1.1-flash vs 1M on the docs page. Exact I/O day not verified.","verified":"2026-09-29","body_md":"Google's \"create anything from any input\" model, starting with video: generate and iteratively edit clips via conversation. Also in Gemini app, Flow, YouTube Shorts, Google Vids.\n\n```python\nfrom google import genai\nclient = genai.Client()\ninteraction = client.interactions.create(model=\"gemini-omni-1.1-flash\",\n    input=\"A corgi surfing a wave at sunset, cinematic, with ocean sounds\")\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [changelog](https://ai.google.dev/gemini-api/docs/changelog), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/). Interactions SDK call shape taken from the Lyria docs; see the video-generation guide for exact output handling.","page_url":"https://postcutoff.com/m/gemini-omni-flash/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":694,"you_url":null,"briefings":null},{"id":"grok-build-0-1","name":"Grok Build 0.1","org":"xAI","family":"Grok Build","released":"2026-05","status":"current","type":"code","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1,"cached_input":0.2,"output":2,"input_over_200k":2,"output_over_200k":4,"unit":"per 1M tokens (USD)","source":"https://docs.x.ai/docs/models/grok-build-0.1"},"price_line":"$1 in, $2 out per 1M tokens","access":[{"provider":"xAI API","model_id":"grok-build-0.1","endpoint":"https://api.x.ai/v1/chat/completions","docs":"https://docs.x.ai/docs/models/grok-build-0.1"},{"provider":"OpenRouter","model_id":"x-ai/grok-build-0.1","url":"https://openrouter.ai/x-ai/grok-build-0.1"}],"capabilities":[{"name":"Agentic coding model","detail":"Reasoning model tuned for agentic software engineering and workflow tasks; powers xAI's Grok Build coding agent.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models/grok-build-0.1"},{"name":"Low-cost coding tier","detail":"$1/$2 per 1M tokens with 256K context - cheapest current Grok text model.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models"}],"entry":null,"notes":"xAI coding model (successor to grok-code-fast line). Release month from OpenRouter listing (2026-05-20); exact date not verified.","verified":"2026-09-29","body_md":"Use for coding agents, IDE integrations and repo-level edits.\n\n```bash\ncurl https://api.x.ai/v1/chat/completions -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-build-0.1\",\"messages\":[{\"role\":\"user\",\"content\":\"Write a Python quicksort\"}]}'\n```\n\nSources: https://docs.x.ai/docs/models/grok-build-0.1 , https://docs.x.ai/docs/models","page_url":"https://postcutoff.com/m/grok-build-0-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":694,"you_url":null,"briefings":null},{"id":"mistral-medium-3-5","name":"Mistral Medium 3.5","org":"Mistral AI","family":"Mistral Medium","released":"2026-04-28","status":"current","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"modified-mit","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1.5,"output":7.5,"unit":"per 1M tokens (USD)","source":"https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04"},"price_line":"$1.50 in, $7.50 out per 1M tokens","access":[{"provider":"Mistral API","model_id":"mistral-medium-3-5","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04"},{"provider":"OpenRouter","model_id":"mistralai/mistral-medium-3-5","url":"https://openrouter.ai/mistralai/mistral-medium-3-5"},{"provider":"Hugging Face","url":"https://huggingface.co/mistralai/Mistral-Medium-3.5-128B"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"One model replacing Devstral 2 and Magistral","detail":"Frontier-class multimodal model for agentic and coding use; Mistral names it the replacement for deprecated Devstral 2 (deprecated 2026-05-22).","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/devstral-2-25-12"},{"name":"Open-weight 128B dense with vision","detail":"128B dense weights on Hugging Face under a modified MIT license, 256K context, built-in tools and Agents/Conversations API support.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04"}],"entry":null,"notes":"Alias mistral-medium-latest (version v26.04). Official card lists 2 more aliases not verified. Batch API supported (OpenRouter batch $0.75/$3.75). Knowledge cutoff not published.","verified":"2026-09-29","body_md":"Mistral's current frontier model for coding agents, reasoning and vision.\n\n```bash\ncurl https://api.mistral.ai/v1/chat/completions -H \"Authorization: Bearer $MISTRAL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"mistral-medium-3-5\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.mistral.ai/models/overview , https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04","page_url":"https://postcutoff.com/m/mistral-medium-3-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":726,"you_url":"https://postcutoff.com/you/mistral-medium-3-5/","briefings":null},{"id":"nemotron-3-nano-omni","name":"NVIDIA Nemotron 3 Nano Omni (30B-A3B Reasoning)","org":"NVIDIA","family":"Nemotron 3","released":"2026-04-28","status":"current","type":"multimodal","modality_in":["text","image","audio","video"],"modality_out":["text"],"open_weights":true,"model_license":"nvidia-open-model-agreement","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"NVIDIA API (build.nvidia.com)","model_id":"nvidia/nemotron-3-nano-omni-30b-a3b-reasoning","endpoint":"https://integrate.api.nvidia.com/v1/chat/completions","docs":"https://build.nvidia.com"},{"provider":"OpenRouter","model_id":"nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free","url":"https://openrouter.ai/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning"},{"provider":"Hugging Face","url":"https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16"}],"capabilities":[{"name":"Open omni-modal reasoning (video + audio + image)","detail":"Single 3B-active open model reasoning over video (up to ~2 min), audio, images and text with chain-of-thought on by default.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16"},{"name":"ASR with word timestamps, OCR, GUI automation","detail":"Targets transcription with word-level timestamps, document intelligence/OCR and GUI agent workflows.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16"}],"entry":null,"notes":"Also FP8/NVFP4 repos. Only free OpenRouter variant seen; paid pricing not verified. Knowledge cutoff not published.","verified":"2026-09-29","body_md":"Small open model for video/speech analysis, document intelligence and GUI agents.\n\n```bash\ncurl https://integrate.api.nvidia.com/v1/chat/completions -H \"Authorization: Bearer $NVIDIA_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"nvidia/nemotron-3-nano-omni-30b-a3b-reasoning\",\"messages\":[{\"role\":\"user\",\"content\":\"Describe this image\"}]}'\n```\n\nSources: https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 , https://integrate.api.nvidia.com/v1/models","page_url":"https://postcutoff.com/m/nemotron-3-nano-omni/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":726,"you_url":"https://postcutoff.com/you/nemotron-3-nano-omni/","briefings":null},{"id":"deepseek-v4-pro","name":"DeepSeek-V4-Pro","org":"DeepSeek","family":"DeepSeek V4","released":"2026-04-24","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":1000000,"max_output":384000,"knowledge_cutoff":null,"pricing":{"input":1.32,"output":3.96,"cache_read":0.044,"unit":"per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.66, output 1.98, cache hit 0.022). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri","source":"https://api-docs.deepseek.com/quick_start/pricing"},"price_line":"$1.32 in, $3.96 out per 1M tokens","access":[{"provider":"DeepSeek API","model_id":"deepseek-v4-pro","endpoint":"https://api.deepseek.com","docs":"https://api-docs.deepseek.com/quick_start/pricing"},{"provider":"DeepSeek API (Anthropic format)","model_id":"deepseek-v4-pro","endpoint":"https://api.deepseek.com/anthropic","docs":"https://api-docs.deepseek.com/quick_start/pricing"},{"provider":"Alibaba Cloud Model Studio","model_id":"deepseek-v4-pro-0813","docs":"https://www.alibabacloud.com/help/en/model-studio/models"},{"provider":"OpenRouter","model_id":"deepseek/deepseek-v4-pro-0813","url":"https://openrouter.ai/deepseek/deepseek-v4-pro-0813"},{"provider":"OpenRouter (preview 0423)","model_id":"deepseek/deepseek-v4-pro","url":"https://openrouter.ai/deepseek/deepseek-v4-pro"},{"provider":"Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813"},{"provider":"Web app","url":"https://chat.deepseek.com"}],"capabilities":[{"name":"Open-weight 1.6T MoE with 1M context","detail":"1.6T total / 49B active parameters, MIT license, 1M-token context (paper: 'Towards Highly Efficient Million-Token Context Intelligence').","first":false,"discovered":"launch","source":"https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash"},{"name":"Agentic GA upgrade (0813)","detail":"GA release greatly strengthened agent performance in production (e.g. Terminal Bench 2.1 87.9, Toolathlon-Verified 74.1 per DeepSeek).","first":false,"discovered":"later","source":"https://api-docs.deepseek.com/updates"},{"name":"Reasoning effort low/high/max","detail":"Thinking mode supports three effort levels; non-thinking mode also available.","first":false,"discovered":"later","source":"https://api-docs.deepseek.com/updates"},{"name":"Native OpenAI Responses API + Codex","detail":"DeepSeek API natively speaks the Responses API format and is adapted for Codex; Anthropic Messages format also supported.","first":false,"discovered":"later","source":"https://api-docs.deepseek.com/updates"},{"name":"DSpark speculative decoding module","detail":"0813 weights ship with an attached DSpark speculative-decoding module for faster inference.","first":false,"discovered":"later","source":"https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813"}],"entry":null,"notes":"Preview 2026-04-24, GA snapshot DeepSeek-V4-Pro-0813 on 2026-08-13 (same id deepseek-v4-pro). Text-only (no vision). DeepSeek said service continues past 2026-09-14 until further notice. Knowledge cutoff not published.","verified":"2026-09-29","body_md":"DeepSeek's strongest model: long-horizon coding agents, reasoning, 1M-context tasks. Open weights (MIT).\n\n```bash\ncurl https://api.deepseek.com/chat/completions \\\n -H \"Authorization: Bearer $DEEPSEEK_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"deepseek-v4-pro\",\"reasoning_effort\":\"high\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://api-docs.deepseek.com/quick_start/pricing · https://api-docs.deepseek.com/updates · https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro-0813","page_url":"https://postcutoff.com/m/deepseek-v4-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":729,"you_url":"https://postcutoff.com/you/deepseek-v4-pro/","briefings":null},{"id":"gpt-5-5","name":"GPT-5.5","org":"OpenAI","family":"GPT-5","released":"2026-04-24","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2025-12","pricing":{"input":5,"cached_input":0.5,"output":30,"unit_note":"prompts >272K input tokens billed 2x input / 1.5x output","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$5 in, $30 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.5","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.5"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.5","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.5","url":"https://openrouter.ai/openai/gpt-5.5"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"xhigh reasoning effort","detail":"Reasoning effort none/low/medium/high/xhigh.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.5"},{"name":"1M context","detail":"1.05M context window with 128K output.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.5"}],"entry":null,"notes":"Snapshot gpt-5.5-2026-04-23. Superseded by GPT-5.6 and GPT-6; still available.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.5\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.5\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-5/","events_after":809,"major_after":202,"historic_after":40,"missing_at_launch":79,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-12-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-12-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-12-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-12.md"}},{"id":"gpt-5-5-pro","name":"GPT-5.5 Pro","org":"OpenAI","family":"GPT-5","released":"2026-04-24","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2025-12","pricing":{"input":30,"output":180,"unit":"per 1M tokens (USD), no cached-input discount","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$30 in, $180 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.5-pro","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.5-pro"},{"provider":"OpenRouter","model_id":"openai/gpt-5.5-pro","url":"https://openrouter.ai/openai/gpt-5.5-pro"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Extended compute","detail":"Uses more compute per request; some requests take several minutes. Effort medium/high/xhigh.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.5-pro"},{"name":"Responses/Batch only","detail":"Not available on Chat Completions.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.5-pro"}],"entry":null,"notes":"Last separately-billed \"-pro\" API id; for GPT-5.6/GPT-6 OpenRouter exposes pro as reasoning.mode pro. Snapshot gpt-5.5-pro-2026-04-23. Azure id not verified.","verified":"2026-09-29","body_md":"Hardest problems where latency does not matter.\n\n```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.5-pro\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.5-pro\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-5-pro/","events_after":809,"major_after":202,"historic_after":40,"missing_at_launch":79,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-5-pro/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-12-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-12-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-12-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-12.md"}},{"id":"gemini-embedding-2","name":"Gemini Embedding 2","org":"Google DeepMind","family":"Gemini Embedding","released":"2026-04-22","status":"current","type":"embedding","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":8192,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.2,"unit":"per 1M text input tokens; image $0.45/1M (~$0.00012 per image), audio $6.50/1M (~$0.00016/s), video $12.00/1M (~$0.00079 per frame)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.20 in per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-embedding-2","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-embedding-2:embedContent","docs":"https://ai.google.dev/gemini-api/docs/embeddings"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/embedding-2"}],"capabilities":[{"name":"Natively multimodal embeddings","detail":"Text, images, video, audio and PDFs mapped into one embedding space; Google's first natively multimodal embedding model and first in the Gemini API.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/embeddings"},{"name":"Matryoshka dimensions","detail":"Flexible 128-3072 output dimensions (recommended 768/1536/3072); 100+ languages.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/embeddings"}],"entry":null,"notes":"Public preview March 2026 (id gemini-embedding-2-preview, still listed), GA 2026-04-22. modality_out 'text' is a placeholder: output is a vector. Predecessor gemini-embedding-001 (text-only) shuts down 2028-05-14.","verified":"2026-09-29","body_md":"Embeddings for multimodal RAG, semantic search, clustering and classification.\n\n```python\nfrom google import genai\nclient = genai.Client()\nr = client.models.embed_content(model=\"gemini-embedding-2\", contents=\"What is the meaning of life?\")\n```\n\nSources: [embeddings guide](https://ai.google.dev/gemini-api/docs/embeddings), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-embedding-2/).","page_url":"https://postcutoff.com/m/gemini-embedding-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":731,"you_url":null,"briefings":null},{"id":"gpt-image-2","name":"GPT Image 2","org":"OpenAI","family":"GPT Image","released":"2026-04-21","status":"legacy","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"text_input":5,"text_cached_input":1.25,"image_input":8,"image_cached_input":2,"image_output":30,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$5 text in, $30 image out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-image-2","endpoint":"https://api.openai.com/v1/images/generations","docs":"https://developers.openai.com/api/docs/models/gpt-image-2"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-image-2","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Batch image generation","detail":"Supports v1/batch in addition to generations/edits.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-image-2"},{"name":"DALL-E replacement","detail":"Named replacement for dall-e-2/dall-e-3 (shut down May 12 2026).","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":null,"notes":"Snapshot gpt-image-2-2026-04-21. Superseded by GPT Image 2.5 Sunburst/Flare; still priced and not deprecated. gpt-image-1 ($10/$40 image) also still listed.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/images/generations \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-image-2\",\"prompt\":\"a lighthouse at dawn\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-image-2\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-image-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":732,"you_url":null,"briefings":null},{"id":"gpt-rosalind","name":"GPT-Rosalind","org":"OpenAI","family":"GPT-Rosalind","released":"2026-04-17","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":5,"cached_input":0.5,"output":25,"unit":"per 1M tokens (USD); billed from 2026-10-05","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"OpenAI API (trusted access only)","model_id":"gpt-rosalind-research","endpoint":"https://api.openai.com/v1/chat/completions","docs":"https://help.openai.com/en/articles/20001193-introducing-gpt-rosalind-for-life-sciences-research"},{"provider":"ChatGPT / Codex (eligible organisations)","url":"https://openai.com/gpt-rosalind/"}],"capabilities":[{"name":"Life-sciences specialist reasoning","detail":"Tuned for genomics, protein and sequence analysis, medicinal chemistry, literature synthesis, wet-lab troubleshooting and experiment planning; OpenAI reports BixBench pass@1 0.751 at launch and LabWorkBench 63.2% (vs GPT-5.5 55.8%) after the June update.","first":false,"discovered":"launch","source":"https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind/"},{"name":"Trusted-access dual-use deployment","detail":"Callable only by vetted organisations with an approved research deployment; a Rosalind Biodefense programme extends access to US government and allied public-health partners.","first":false,"discovered":"launch","source":"https://www.rdworldonline.com/openai-launches-rosalind-biodefense-offers-federal-agencies-early-access-to-its-life-sciences-model/"}],"entry":"2026-04-17-openai-gpt-rosalind","notes":"Research preview 17 Apr 2026; rebuilt on GPT-5.5 on 3 June 2026 (OpenAI says 31% fewer tokens than GPT-5.5); out of preview globally 11 Sept 2026. Context window and max output not published. Pricing per OpenAI's pricing page 'Life Sciences' section, as quoted by TokenCost and the Portkey model registry (PR #953); not read directly on openai.com (403). Verified 2026-10-05 on developers.openai.com/api/docs/pricing ('Life Sciences' row: $5.00 / $0.50 cached / $25.00; 'Billing for gpt-rosalind-research begins on October 5, 2026. Cache-write pricing does not apply'; access limited to approved internal research through the trusted-access program). Free Codex Life Sciences plugin connects any model to 50+ scientific tools.","verified":"2026-10-05","body_md":"Access requires OpenAI trusted-access approval; ordinary API keys will be refused.\n\nSources:\n- https://openai.com/index/introducing-gpt-rosalind/\n- https://github.com/Portkey-AI/models/pull/953","page_url":"https://postcutoff.com/m/gpt-rosalind/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":733,"you_url":"https://postcutoff.com/you/gpt-rosalind/","briefings":null},{"id":"grok-tts","name":"Grok Text to Speech (Grok TTS API)","org":"xAI","family":"Grok Voice","released":"2026-04-17","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_million_characters":15,"unit":"USD per 1M characters","source":"https://docs.x.ai/developers/pricing"},"price_line":"$15 per 1M characters","access":[{"provider":"xAI API (REST)","endpoint":"https://api.x.ai/v1/tts","docs":"https://docs.x.ai/developers/model-capabilities/audio/text-to-speech"},{"provider":"xAI API (WebSocket streaming)","endpoint":"wss://api.x.ai/v1/tts","docs":"https://docs.x.ai/developers/model-capabilities/audio/text-to-speech"}],"capabilities":[{"name":"Inline speech tags","detail":"Inline tags ([pause], [laugh], [sigh], [cry], [gasp], ...) and wrapping tags (<whisper>, <soft>, <loud>, <slow>, <fast>, <sing>) control delivery.","first":false,"discovered":"launch","source":"https://docs.x.ai/developers/model-capabilities/audio/text-to-speech"},{"name":"Custom (cloned) voices","detail":"Clone a voice from a short reference clip via the Custom Voices API; the voice_id works like built-in voices in TTS and the Voice Agent API.","first":false,"discovered":"later","source":"https://docs.x.ai/developers/model-capabilities/audio/text-to-speech"}],"entry":null,"notes":"Launched with the Grok STT API on 2026-04-17 (some press reports an earlier developer opening in March 2026). No separate model id is documented; the endpoint selects the model. 60,000 characters per REST request; ~20 languages plus auto-detect; MP3/WAV/PCM/mu-law/A-law at 8-48 kHz; voice list via GET /v1/tts/voices (Ara, Eve, Leo, Rex, Sal and many more).","verified":"2026-09-29","body_md":"Call `POST https://api.x.ai/v1/tts` with your text and a `voice` (see the docs for the exact request schema).\n\nSources: https://docs.x.ai/developers/model-capabilities/audio/text-to-speech · https://x.ai/news/grok-stt-and-tts-apis · https://docs.x.ai/developers/pricing","page_url":"https://postcutoff.com/m/grok-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":733,"you_url":null,"briefings":null},{"id":"gr00t-n1-7","name":"Isaac GR00T N1.7","org":"NVIDIA","family":"Isaac GR00T","released":"2026-04-17","status":"current","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":true,"model_license":"NVIDIA Open Model License (weights, commercial use); Apache-2.0 (code)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/GR00T-N1.7-3B","url":"https://huggingface.co/nvidia/GR00T-N1.7-3B","docs":"https://huggingface.co/blog/nvidia/gr00t-n1-7"},{"provider":"GitHub","url":"https://github.com/NVIDIA/Isaac-GR00T"},{"provider":"Hugging Face LeRobot integration","docs":"https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/"}],"capabilities":[{"name":"Human egocentric video pretraining","detail":"Pretrained on 20,854 hours of human egocentric video (EgoScale) across 20+ task categories, on top of robot data.","first":false,"discovered":"launch","source":"https://huggingface.co/blog/nvidia/gr00t-n1-7"},{"name":"Scaling law for robot dexterity","detail":"NVIDIA reports the 'first-ever scaling law for robot dexterity': more human video predictably improves 22-DoF hand performance without mass teleoperation.","first":true,"discovered":"launch","source":"https://huggingface.co/blog/nvidia/gr00t-n1-7"},{"name":"Reasoning VLA on a Cosmos backbone","detail":"3B 'Action Cascade' model: Cosmos-Reason2-2B VLM plus 32-layer diffusion transformer; relative end-effector action space; runs on one 16 GB+ GPU including Jetson Thor/Orin and DGX Spark.","first":false,"discovered":"launch","source":"https://github.com/NVIDIA/Isaac-GR00T"}],"entry":"2026-03-16-nvidia-gtc-2026-robotics-groot-n2-cosmos-3","notes":"Early access with commercial licensing announced at GTC 2026-03-16; open release/HF blog 2026-04-17. Post-trained checkpoints: GR00T-N1.7-LIBERO, -DROID, -SimplerEnv-Bridge, -SimplerEnv-Fractal, GR00T-H-N1.7 (surgical-robotics variant, uploaded to HF 2026-05-30: 3B, post-trained on 601 h / ~63.9k episodes of real surgical tasks from the Open-H-Embodiment dataset across 7 platforms incl. dVRK, CMR Versius, KUKA LBR iiwa; NVIDIA Open Model License; R&D only, not for clinical use; follows the original GR00T-H announced at GTC 2026-03-16). Backbone nvidia/Cosmos-Reason2-2B is gated (accept license on HF). Validated on Unitree G1, YAM bimanual, AGIBot Genie 1. Fine-tuning: 40 GB+ GPUs recommended.","verified":"2026-09-29","body_md":"The current open GR00T release; GR00T N2 (world action model) is due by end of 2026.\n\nChangelog: 2026-09-29 added GR00T-H-N1.7 details.\n\nSources: [GR00T-H-N1.7 model card](https://huggingface.co/nvidia/GR00T-H-N1.7), [HF blog](https://huggingface.co/blog/nvidia/gr00t-n1-7), [GitHub](https://github.com/NVIDIA/Isaac-GR00T), [NVIDIA newsroom (GTC 2026)](https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world).","page_url":"https://postcutoff.com/m/gr00t-n1-7/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":733,"you_url":null,"briefings":null},{"id":"claude-opus-4-7","name":"Claude Opus 4.7","org":"Anthropic","family":"Claude 4","released":"2026-04-16","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2026-01","pricing":{"input":5,"output":25,"cache_read":0.5,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-4-7","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-4-7/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-opus-4-7","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-4-7","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-4-7","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-4-7","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-4.7","url":"https://openrouter.ai/anthropic/claude-opus-4.7"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"High-resolution vision","detail":"Accepts images up to 2,576 px on the long edge (~3.75 MP), more than 3x prior Claude models. Scored 98.5% on XBOW visual acuity versus 54.5% for Opus 4.6.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-7"},{"name":"xhigh effort level","detail":"Introduced the xhigh effort level between high and max.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-7"},{"name":"New tokenizer","detail":"First model with Anthropic's newer tokenizer (about 30% more tokens for the same text).","first":true,"discovered":"launch","source":"https://platform.claude.com/docs/en/about-claude/pricing"},{"name":"Hard coding tasks","detail":"Resolved about 3x more production tasks than Opus 4.6 on Rakuten-SWE-Bench.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-7"}],"entry":null,"notes":"Introduced the newer tokenizer (~30% more tokens per text) and xhigh effort. Adaptive thinking only. Retirement not sooner than 2027-04-16.","verified":"2026-09-29","body_md":"Legacy Opus; migrate to `claude-opus-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-4-7\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-4-7/overview\n- Announcement: https://www.anthropic.com/news/claude-opus-4-7\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-4-7/","events_after":793,"major_after":197,"historic_after":40,"missing_at_launch":56,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-4-7/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2026-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2026-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2026-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2026-01.md"}},{"id":"pi-0-7","name":"π0.7","org":"Physical Intelligence","family":"π (pi)","released":"2026-04-16","status":"current","type":"robotics","modality_in":["text","image"],"modality_out":["action","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"None (internal / partner deployments; no public weights or API)","url":"https://www.pi.website/blog/pi07"}],"capabilities":[{"name":"Compositional generalization to untrained tasks","detail":"Recombines skills to do tasks never in training (e.g. operating an air fryer seen only in two fragmentary training episodes; laundry folding on a robot with no folding data).","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pi07"},{"name":"Steerable by natural-language coaching","detail":"Plain-language coaching lifted air-fryer success from ~5% to ~95% in about 30 minutes, without retraining.","first":false,"discovered":"launch","source":"https://techcrunch.com/2026/04/16/physical-intelligence-a-hot-robotics-startup-says-its-new-robot-brain-can-figure-out-tasks-it-was-never-taught/"},{"name":"Generalist matches fine-tuned specialists","detail":"One general model performs dexterous tasks at the level of per-task fine-tuned specialists and transfers across embodiments.","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pi07"}],"entry":"2026-04-16-physical-intelligence-pi-0-7","notes":"PI describes 'the first signs of compositional generalization' in its own models; not marked first:true. No weights in openpi as of 2026-09-29 (latest open PI model is π0.5). Parameter count not found. No newer PI model found through 2026-09-29.","verified":"2026-09-29","body_md":"Sources: [π0.7 blog](https://www.pi.website/blog/pi07), [TechCrunch](https://techcrunch.com/2026/04/16/physical-intelligence-a-hot-robotics-startup-says-its-new-robot-brain-can-figure-out-tasks-it-was-never-taught/).","page_url":"https://postcutoff.com/m/pi-0-7/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":735,"you_url":null,"briefings":null},{"id":"gemini-3-1-flash-tts","name":"Gemini 3.1 Flash TTS (preview)","org":"Google DeepMind","family":"Gemini 3.1","released":"2026-04-15","status":"legacy","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1,"output":20,"unit":"per 1M tokens (USD), text in / audio out (25 audio tokens per second)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1 in, $20 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.1-flash-tts-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-tts-preview:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"Google Cloud Text-to-Speech (Gemini-TTS)","model_id":"Gemini 3.1 Flash TTS (Preview)","docs":"https://cloud.google.com/text-to-speech/pricing"}],"capabilities":[{"name":"Steerable expressive TTS","detail":"'Cost-efficient, expressive, and steerable text to speech' controlled with natural-language prompts.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/changelog"}],"entry":null,"notes":"Superseded by gemini-3.8-flash-tts (GA 2026-09-22), which is cheaper at intro pricing ($0.50/$9.00). Still served as preview; also billed in Cloud TTS at the same $1/$20.","verified":"2026-09-29","body_md":"Sources: [models](https://ai.google.dev/gemini-api/docs/models), [changelog](https://ai.google.dev/gemini-api/docs/changelog), [Gemini pricing](https://ai.google.dev/gemini-api/docs/pricing), [Cloud TTS pricing](https://cloud.google.com/text-to-speech/pricing).","page_url":"https://postcutoff.com/m/gemini-3-1-flash-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":737,"you_url":null,"briefings":null},{"id":"agibot-go-2","name":"AgiBot GO-2 (Genie Operator-2)","org":"AgiBot","family":"Genie Operator","released":"2026-04-09","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"AgiBot robots / Genie Studio (via AgiBot sales)","url":"https://www.agibot.com/article/231/detail/56.html"}],"capabilities":[{"name":"Action chain-of-thought","detail":"Reasons in action space: generates a macro-plan of high-level action intents, then executes step by step, with teacher forcing so execution adheres to the reasoning.","first":false,"discovered":"launch","source":"https://www.agibot.com/article/231/detail/56.html"},{"name":"Asynchronous dual-system","detail":"Low-frequency semantic planner ('commander') plus high-frequency action follower ('executor') in one architecture.","first":false,"discovered":"launch","source":"https://www.therobotreport.com/agibot-releases-go-2-foundation-model-embodied-ai/"},{"name":"Benchmark results","detail":"LIBERO 98.5% average (ranked 1st), LIBERO-Plus 86.6% zero-shot, VLABench 47.4, 82.9% real-world success from simulation-only training (company-reported).","first":false,"discovered":"launch","source":"https://www.agibot.com/article/231/detail/56.html"}],"entry":"2026-04-09-agibot-go-2","notes":"No open weights, API or pricing found (GO-1 was open, non-commercial). Core work accepted to CVPR 2026 and ACL 2026 per AgiBot. Trained on 'tens of thousands of hours' of interaction data.","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/agibot-go-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":740,"you_url":null,"briefings":null},{"id":"gemma-4","name":"Gemma 4","org":"Google DeepMind","family":"Gemma","released":"2026-04-02","status":"current","type":"llm","modality_in":["text","image","audio","video"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.09,"output":0.34,"unit":"per 1M tokens, OpenRouter price for google/gemma-4-31b-it (26B-A4B: $0.09 / $0.30; free variants exist). Weights free to download","source":"https://openrouter.ai/google/gemma-4-31b-it"},"price_line":"$0.09 in, $0.34 out per 1M tokens","access":[{"provider":"Hugging Face","model_id":"google/gemma-4-31B-it","url":"https://huggingface.co/google/gemma-4-31B-it"},{"provider":"Hugging Face","model_id":"google/gemma-4-26B-A4B-it","url":"https://huggingface.co/google/gemma-4-26B-A4B-it"},{"provider":"Hugging Face","model_id":"google/gemma-4-12B-it","url":"https://huggingface.co/google/gemma-4-12B-it"},{"provider":"Hugging Face","model_id":"google/gemma-4-E4B-it","url":"https://huggingface.co/google/gemma-4-E4B-it"},{"provider":"Hugging Face","model_id":"google/gemma-4-E2B-it","url":"https://huggingface.co/google/gemma-4-E2B-it"},{"provider":"OpenRouter","model_id":"google/gemma-4-31b-it"},{"provider":"OpenRouter","model_id":"google/gemma-4-26b-a4b-it"},{"provider":"Google docs","docs":"https://ai.google.dev/gemma/docs/core"}],"capabilities":[{"name":"First Apache-2.0 Gemma","detail":"First Gemma generation under the permissive Apache 2.0 license instead of Google's custom Gemma terms.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/"},{"name":"Intelligence per parameter","detail":"31B dense ranked #3 and 26B A4B MoE #6 among open models on Arena at launch.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/"},{"name":"On-device agentic models","detail":"E2B/E4B edge models with native audio+vision, function calling and structured JSON, running offline on phones/Raspberry Pi/Jetson.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/"},{"name":"Encoder-free unified 12B","detail":"Gemma 4 12B, added later, is a unified encoder-free multimodal model with native audio.","first":false,"discovered":"later","source":"https://ai.google.dev/gemma/docs/core"}],"entry":null,"notes":"Sizes E2B, E4B (128K context), 12B, 26B A4B MoE, 31B dense (256K context); base and -it variants plus QAT/GGUF quantized repos. 12B released later (HF repo 2026-05-23). Audio input on E2B/E4B/12B only. 140+ languages. pricing is third-party (OpenRouter), not Google.","verified":"2026-09-29","body_md":"Latest open-weights family from Google DeepMind (built from Gemini 3 research); run locally or self-host.\n\n```python\nfrom transformers import pipeline\npipe = pipeline(\"image-text-to-text\", model=\"google/gemma-4-E4B-it\")\nprint(pipe(text=[{\"role\":\"user\",\"content\":[{\"type\":\"text\",\"text\":\"Hello!\"}]}]))\n```\n\nSources: [Gemma docs](https://ai.google.dev/gemma/docs/core), [launch blog](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/), [HF repo](https://huggingface.co/google/gemma-4-31B-it).","page_url":"https://postcutoff.com/m/gemma-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":745,"you_url":"https://postcutoff.com/you/gemma-4/","briefings":null},{"id":"generalist-gen-1","name":"Generalist GEN-1","org":"Generalist AI","family":"GEN","released":"2026-04-02","status":"current","type":"robotics","modality_in":["image","video","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Generalist AI early-access partners (partnerships@generalistai.com)","url":"https://generalistai.com/blog/gen-1"}],"capabilities":[{"name":"Mastery of simple physical tasks","detail":"99% success on several tasks (GEN-0: 64%), ~3x faster than prior state of the art, ~1 hour of robot data per task; Generalist calls it the first general-purpose model to cross a 'mastery' threshold for simple tasks.","first":true,"discovered":"launch","source":"https://generalistai.com/blog/gen-1"},{"name":"Pretrained on 500k+ hours of human wearable data","detail":"Pretraining dataset of 500,000+ hours of real-world physical interaction captured with wearable devices on humans (no robot data), spanning many end effectors; later extended to a broad range of end effectors from five-finger hands to special tools.","first":false,"discovered":"launch","source":"https://generalistai.com/blog/gen-1"},{"name":"Robotics scaling laws (GEN-0 predecessor)","detail":"GEN-0 (Nov 2025) showed scaling laws for robot foundation models, with all tracked zero-shot tasks improving together as pretraining scaled.","first":false,"discovered":"launch","source":"https://generalistai.com/blog/gen-1"}],"entry":"2026-04-02-generalist-gen-1","notes":"No public API/weights; early-access partners only. Successor GEN-1.5 (2026-08-19) adds one-shot learning (see generalist-gen-1-5). Results are company-reported. Video: https://www.youtube.com/watch?v=SY2xyrmV44Y","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/generalist-gen-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":745,"you_url":null,"briefings":null},{"id":"grok-4-3","name":"Grok 4.3","org":"xAI","family":"Grok 4","released":"2026-04","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1.25,"cached_input":0.2,"output":2.5,"input_over_200k":2.5,"output_over_200k":5,"unit":"per 1M tokens (USD); Batch API 20% off","source":"https://docs.x.ai/docs/models"},"price_line":"$1.25 in, $2.50 out per 1M tokens","access":[{"provider":"xAI API","model_id":"grok-4.3","endpoint":"https://api.x.ai/v1/chat/completions","docs":"https://docs.x.ai/docs/models/grok-4.3"},{"provider":"AWS Bedrock","model_id":"xai.grok-4.3","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-3.html"},{"provider":"OpenRouter","model_id":"x-ai/grok-4.3","url":"https://openrouter.ai/x-ai/grok-4.3"},{"provider":"Web app","url":"https://grok.com"}],"capabilities":[{"name":"1M context at budget price","detail":"1M-token context window at $1.25/$2.50, cheaper than the 500K-context Grok 4.5-4.7 line.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models/grok-4.3"},{"name":"Reasoning effort incl. none","detail":"Reasoning effort none / low / medium / high / xhigh, default low - usable as a fast non-reasoning model.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models/grok-4.3"}],"entry":null,"notes":"Alias grok-4.3-latest. Cheaper long-context option still offered alongside Grok 4.7. Release month inferred from OpenRouter listing date (2026-04-30); exact date not verified.","verified":"2026-09-29","body_md":"Cost-efficient 1M-context Grok for long documents and high-volume agents.\n\n```bash\ncurl https://api.x.ai/v1/chat/completions -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-4.3\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.x.ai/docs/models , https://docs.x.ai/docs/models/grok-4.3","page_url":"https://postcutoff.com/m/grok-4-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":725,"you_url":"https://postcutoff.com/you/grok-4-3/","briefings":null},{"id":"kimi-k2-6","name":"Kimi K2.6","org":"Moonshot AI","family":"Kimi K2","released":"2026-04","status":"current","type":"multimodal","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"modified-mit","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.95,"output":4,"cache_read":0.16,"unit":"per 1M tokens (USD)","source":"https://platform.kimi.ai/docs/pricing/chat"},"price_line":"$0.95 in, $4 out per 1M tokens","access":[{"provider":"Kimi API (Moonshot)","model_id":"kimi-k2.6","endpoint":"https://api.moonshot.ai/v1","docs":"https://platform.kimi.ai/docs/guide/kimi-k2-6-quickstart"},{"provider":"Alibaba Cloud Model Studio","model_id":"kimi-k2.6","docs":"https://www.alibabacloud.com/help/en/model-studio/text-generation-model"},{"provider":"OpenRouter","model_id":"moonshotai/kimi-k2.6","url":"https://openrouter.ai/moonshotai/kimi-k2.6"},{"provider":"Hugging Face","url":"https://huggingface.co/moonshotai/Kimi-K2.6"},{"provider":"Web app","url":"https://www.kimi.com"}],"capabilities":[{"name":"Native multimodal open agentic model","detail":"1T total / 32B active MoE with 400M vision encoder; text, image and video input; thinking and non-thinking modes.","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2.6"},{"name":"Swarm-based task orchestration","detail":"Marketed for proactive autonomous execution and agent-swarm orchestration plus coding-driven design.","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2.6"}],"entry":null,"notes":"Still offered on the API alongside K3 (only remaining non-coding K2-series model; kimi-k2.5 discontinued 2026-08-31). Release day not verified (HF 2026-04-14, OpenRouter 2026-04-20).","verified":"2026-09-29","body_md":"Cheaper multimodal Kimi for chat, visual understanding and agent tasks with switchable thinking.\n\n```bash\ncurl https://api.moonshot.ai/v1/chat/completions \\\n -H \"Authorization: Bearer $MOONSHOT_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"kimi-k2.6\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://platform.kimi.ai/docs/models · https://platform.kimi.ai/docs/pricing/chat · https://huggingface.co/moonshotai/Kimi-K2.6","page_url":"https://postcutoff.com/m/kimi-k2-6/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":725,"you_url":"https://postcutoff.com/you/kimi-k2-6/","briefings":null},{"id":"gemini-3-1-flash-live","name":"Gemini 3.1 Flash Live (preview)","org":"Google DeepMind","family":"Gemini 3.1","released":"2026-03-26","status":"legacy","type":"audio/speech","modality_in":["text","audio","image","video"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.75,"output":4.5,"audio_input":3,"audio_output":12,"unit":"per 1M tokens (USD); same price as gemini-3.8-live","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.75 in, $4.50 out per 1M tokens","access":[{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-3.1-flash-live-preview","docs":"https://ai.google.dev/gemini-api/docs/models"}],"capabilities":[{"name":"Audio-to-audio real-time dialogue","detail":"Native audio model 'designed for real-time dialogue and voice-first AI applications' on the Live API.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/changelog"}],"entry":null,"notes":"Preview id; the models page labels it legacy and recommends gemini-3.8-live (GA 2026-09-15). No shutdown date announced as of 2026-09-29. Context window not re-checked.","verified":"2026-09-29","body_md":"Sources: [models](https://ai.google.dev/gemini-api/docs/models), [changelog](https://ai.google.dev/gemini-api/docs/changelog), [pricing](https://ai.google.dev/gemini-api/docs/pricing).","page_url":"https://postcutoff.com/m/gemini-3-1-flash-live/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":750,"you_url":null,"briefings":null},{"id":"suno-v5-5","name":"Suno v5.5","org":"Suno","family":"Suno v5","released":"2026-03-26","status":"retired","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Web app","url":"https://suno.com"}],"capabilities":[{"name":"Voices (sing with your own voice)","detail":"Record/upload your voice (with verification and privacy controls) and have Suno sing songs in it; Pro/Premier.","first":false,"discovered":"launch","source":"https://about.suno.com/blog/v5-5"},{"name":"Custom Models","detail":"Fine-tune a personal v5.5 on your own catalog (min. 6 tracks, up to 3 models per user); Pro/Premier.","first":false,"discovered":"launch","source":"https://about.suno.com/blog/v5-5"},{"name":"My Taste personalization","detail":"Learns preferred genres/moods and applies them via the Magic Wand; all users.","first":false,"discovered":"launch","source":"https://about.suno.com/blog/v5-5"}],"entry":"2026-03-26-suno-v5-5-voices-custom-models","notes":"No official public API (web/mobile app only; third-party 'Suno APIs' are unofficial). Retired on 2026-09-09 when Suno moved entirely to the v6 family (see suno-v6); Voices and Custom Models features carried over to v6 plans.","verified":"2026-09-29","body_md":"Former Suno flagship (retired 2026-09-09, replaced by [suno-v6](suno-v6.md)). Was the leading consumer song generator (vocals + full arrangement from a prompt/lyrics), with personalization in v5.5. Use via https://suno.com or the mobile apps.\n\nSources: https://about.suno.com/blog/v5-5 , https://suno.com/release-notes , https://suno.com/blog/introducing-v6","page_url":"https://postcutoff.com/m/suno-v5-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":750,"you_url":null,"briefings":null},{"id":"voxtral-tts","name":"Voxtral TTS","org":"Mistral AI","family":"Voxtral","released":"2026-03-23","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"cc-by-nc-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.016,"unit":"USD per 1K characters ($16 per 1M) on Mistral API","source":"https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03"},"price_line":"$0.016 per 1K characters","access":[{"provider":"Mistral API","model_id":"voxtral-tts-2603","endpoint":"https://api.mistral.ai/v1/audio/speech","docs":"https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03"},{"provider":"Hugging Face","model_id":"mistralai/Voxtral-4B-TTS-2603","url":"https://huggingface.co/mistralai/Voxtral-4B-TTS-2603"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"Zero-shot voice cloning from ~3 s","detail":"Clones a voice (accent, fillers, rhythm) from a few seconds of reference audio without a transcript; 68.4% human-preference win rate vs ElevenLabs Flash v2.5 on multilingual cloning (Mistral-reported).","first":false,"discovered":"launch","source":"https://mistral.ai/news/voxtral-tts"},{"name":"Open-weight 4B TTS with low latency","detail":"3.4B decoder + 390M flow-matching acoustic transformer + 300M codec; ~70 ms model latency (~90 ms time-to-first-audio via API), RTF ~9.7x, up to 2 min native generation.","first":false,"discovered":"launch","source":"https://mistral.ai/news/voxtral-tts"}],"entry":"2026-03-23-mistral-voxtral-tts","notes":"Mistral's first TTS model. 9 languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi, Arabic. Weights are CC BY-NC 4.0 (non-commercial); commercial use via API. Docs model-card page shows id voxtral-tts-2603 on the overview (a 'voxtral-mini-tts-2603' alias also appears on the card).","verified":"2026-09-29","body_md":"Call `POST https://api.mistral.ai/v1/audio/speech` with `model: voxtral-tts-2603` (see https://docs.mistral.ai/capabilities/audio/text_to_speech for the request schema and voice options).\n\nSources: https://mistral.ai/news/voxtral-tts · https://docs.mistral.ai/models/overview · https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03","page_url":"https://postcutoff.com/m/voxtral-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":757,"you_url":null,"briefings":null},{"id":"minimax-m2-7","name":"MiniMax-M2.7","org":"MiniMax","family":"MiniMax-M","released":"2026-03-18","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"other (MiniMax license, see HF)","context_window":204800,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":1.2,"cache_read":0.06,"cache_write":0.375,"unit":"per 1M tokens (USD); MiniMax-M2.7-highspeed: 0.6 / 2.4","source":"https://platform.minimax.io/docs/guides/pricing-paygo"},"price_line":"$0.30 in, $1.20 out per 1M tokens","access":[{"provider":"MiniMax API (Anthropic format)","model_id":"MiniMax-M2.7","endpoint":"https://api.minimax.io/anthropic","docs":"https://platform.minimax.io/docs/api-reference/text-anthropic-api"},{"provider":"MiniMax API (OpenAI format)","model_id":"MiniMax-M2.7","endpoint":"https://api.minimax.io/v1","docs":"https://platform.minimax.io/docs/api-reference/text-openai-api"},{"provider":"MiniMax API (fast)","model_id":"MiniMax-M2.7-highspeed","endpoint":"https://api.minimax.io/v1","docs":"https://platform.minimax.io/docs/api-reference/text-openai-api"},{"provider":"OpenRouter","model_id":"minimax/minimax-m2.7","url":"https://openrouter.ai/minimax/minimax-m2.7"},{"provider":"Hugging Face","url":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7"}],"capabilities":[{"name":"Participates in its own evolution","detail":"MiniMax calls it its first model deeply participating in its own development ('recursive self-improvement').","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7"},{"name":"Agent harness building","detail":"Builds complex agent harnesses using Agent Teams, Skills and dynamic tool search; aimed at professional office delivery.","first":false,"discovered":"launch","source":"https://huggingface.co/MiniMaxAI/MiniMax-M2.7"}],"entry":null,"notes":"Text-only predecessor of M3, still a current API model; highspeed variant ~100 tok/s vs ~60. M2.5/M2.1/M2 are legacy (same $0.3/$1.2 price).","verified":"2026-09-29","body_md":"Budget agentic coding model with 204.8K context; good for self-hosting (SGLang) or cheap API use.\n\n```bash\ncurl https://api.minimax.io/v1/chat/completions \\\n -H \"Authorization: Bearer $MINIMAX_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"MiniMax-M2.7\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://huggingface.co/MiniMaxAI/MiniMax-M2.7","page_url":"https://postcutoff.com/m/minimax-m2-7/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":761,"you_url":"https://postcutoff.com/you/minimax-m2-7/","briefings":null},{"id":"gr00t-n2","name":"Isaac GR00T N2","org":"NVIDIA","family":"Isaac GR00T","released":"2026-03-16","status":"preview","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"not yet released","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Not yet available (NVIDIA says end of 2026)","url":"https://developer.nvidia.com/isaac/gr00t"}],"capabilities":[{"name":"World action model (DreamZero)","detail":"Predicts how the scene will evolve (future latent states) before generating the action sequence; succeeds at new tasks in new environments more than twice as often as leading VLAs (NVIDIA).","first":false,"discovered":"launch","source":"https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world"},{"name":"Top of generalist-policy leaderboards","detail":"NVIDIA says it ranks No. 1 on MolmoSpaces and RoboArena for generalist robot policies (as of GTC, March 2026).","first":false,"discovered":"launch","source":"https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world"}],"entry":"2026-03-16-nvidia-gtc-2026-robotics-groot-n2-cosmos-3","notes":"Previewed in Jensen Huang's GTC keynote 2026-03-16; 'released' = preview date. No weights, API or HF repo found as of 2026-09-29. Modalities assumed from the GR00T line; confirm at release.","verified":"2026-09-29","body_md":"Sources: [NVIDIA newsroom](https://nvidianews.nvidia.com/news/nvidia-and-global-robotics-leaders-take-physical-ai-to-the-real-world), [The Decoder](https://the-decoder.com/gtc-2026-nvidia-wants-to-swap-robotics-data-problem-for-a-compute-problem/).","page_url":"https://postcutoff.com/m/gr00t-n2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":762,"you_url":null,"briefings":null},{"id":"mistral-small-4","name":"Mistral Small 4","org":"Mistral AI","family":"Mistral Small","released":"2026-03-16","status":"current","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.15,"output":0.6,"unit":"per 1M tokens (USD)","source":"https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03"},"price_line":"$0.15 in, $0.60 out per 1M tokens","access":[{"provider":"Mistral API","model_id":"mistral-small-2603","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03"},{"provider":"OpenRouter","model_id":"mistralai/mistral-small-2603","url":"https://openrouter.ai/mistralai/mistral-small-2603"},{"provider":"Hugging Face","url":"https://huggingface.co/mistralai/Mistral-Small-4-119B-2603"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"Instruct + reasoning + coding unified","detail":"First Mistral model unifying Magistral (reasoning), Pixtral (multimodal) and Devstral (agentic coding) in one model; reasoning_effort none/high per request.","first":false,"discovered":"launch","source":"https://mistral.ai/news/mistral-small-4/"},{"name":"119B MoE with ~6.5B active","detail":"119B total / 6.5B active parameters, vision input, 256K context at $0.15/$0.6.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03"}],"entry":null,"notes":"Alias mistral-small-latest (v26.03). Announced Mar 16, 2026.","verified":"2026-09-29","body_md":"Cheap, fast open model for most production workloads, with optional reasoning.\n\n```bash\ncurl https://api.mistral.ai/v1/chat/completions -H \"Authorization: Bearer $MISTRAL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"mistral-small-latest\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03 , https://mistral.ai/news/mistral-small-4/","page_url":"https://postcutoff.com/m/mistral-small-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":762,"you_url":"https://postcutoff.com/you/mistral-small-4/","briefings":null},{"id":"nemotron-3-super","name":"NVIDIA Nemotron 3 Super (120B-A12B)","org":"NVIDIA","family":"Nemotron 3","released":"2026-03-11","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"nvidia-open-model-license","context_window":1000000,"max_output":null,"knowledge_cutoff":"2025-06","pricing":{"input":0.08,"output":0.45,"unit":"per 1M tokens (USD) on OpenRouter; NVIDIA hosted pricing not verified","source":"https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b"},"price_line":"$0.08 in, $0.45 out per 1M tokens","access":[{"provider":"NVIDIA API (build.nvidia.com)","model_id":"nvidia/nemotron-3-super-120b-a12b","endpoint":"https://integrate.api.nvidia.com/v1/chat/completions","docs":"https://build.nvidia.com"},{"provider":"AWS Bedrock","model_id":"nvidia.nemotron-super-3-120b","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html"},{"provider":"OpenRouter","model_id":"nvidia/nemotron-3-super-120b-a12b","url":"https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b"},{"provider":"Hugging Face","url":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16"}],"capabilities":[{"name":"Efficient hybrid LatentMoE for agents","detail":"120B total / 12B active hybrid Mamba-2 + MoE + attention, built for high-volume agentic workloads with up to 1M context (256K default).","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16"},{"name":"Managed on AWS Bedrock","detail":"One of the few NVIDIA open models offered as a serverless Bedrock model (nvidia.nemotron-super-3-120b).","first":false,"discovered":"later","source":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html"}],"entry":null,"notes":"Knowledge cutoff = pre-training (Jun 2025); post-training to Feb 2026. Also FP8/NVFP4 repos. Free tier on OpenRouter (:free).","verified":"2026-09-29","body_md":"Mid-size open reasoning model, widely hosted (Bedrock, NIM, OpenRouter).\n\n```bash\ncurl https://integrate.api.nvidia.com/v1/chat/completions -H \"Authorization: Bearer $NVIDIA_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"nvidia/nemotron-3-super-120b-a12b\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 , https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-nvidia-nemotron-super-3-120b.html","page_url":"https://postcutoff.com/m/nemotron-3-super/","events_after":860,"major_after":218,"historic_after":44,"missing_at_launch":96,"events_since_release":null,"you_url":"https://postcutoff.com/you/nemotron-3-super/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-06.md"}},{"id":"fish-audio-s2","name":"Fish Audio S2 Pro / S2.1 Pro","org":"Fish Audio","family":"Fish Audio S2","released":"2026-03-09","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"Fish Audio Research License (S2 Pro weights; non-commercial, commercial license on request)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_utf8_bytes":15,"unit":"USD per 1M UTF-8 bytes (s1, s2-pro, s2.1-pro); s2.1-pro-free $0 under fair use through 2026-11-30; ASR transcribe-1 $0.36/hr","source":"https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits"},"price_line":"$15 per 1M UTF-8 bytes","access":[{"provider":"Fish Audio API","model_id":"s2.1-pro","endpoint":"https://api.fish.audio/v1/tts (model passed in `model` header)","docs":"https://docs.fish.audio/developer-guide/getting-started/changelog"},{"provider":"Fish Audio API (free tier)","model_id":"s2.1-pro-free","endpoint":"https://api.fish.audio/v1/tts"},{"provider":"Fish Audio API","model_id":"s2-pro"},{"provider":"Hugging Face","model_id":"fishaudio/s2-pro","url":"https://huggingface.co/fishaudio/s2-pro"},{"provider":"OpenRouter","url":"https://openrouter.ai/fish-audio/s2.1-pro"},{"provider":"GitHub","url":"https://github.com/fishaudio/fish-speech"}],"capabilities":[{"name":"Inline natural-language emotion/paralinguistic tags","detail":"Free-form bracket cues like [whisper], [laugh], [emphasis]; multi-speaker dialogue in one pass; 80+ languages from 10M+ hours of training audio.","first":false,"discovered":"launch","source":"https://fish.audio/blog/fish-audio-open-sources-s2/"},{"name":"Open model with production inference stack","detail":"Dual-AR (4B slow + 400M fast) on a Qwen3-4B backbone released with fine-tuning code and SGLang serving; RTF 0.195, ~100 ms TTFA; Seed-TTS Eval WER 0.54% zh / 0.99% en.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2603.08823"},{"name":"Free production API (S2.1 Pro)","detail":"S2.1 Pro (closed, 2026-06-23) offered free under fair use with ~90 ms TTFA, 83 languages; 61% win rate vs S2 Pro.","first":false,"discovered":"later","source":"https://fish.audio/blog/s2-1-pro-free-api/"}],"entry":"2026-03-09-fish-audio-s2-open-source","notes":"S2 Pro held #1 open-weights on Artificial Analysis until Breeze TTS 2 (Aug 2026); now #2 open (~1119 Elo). S2.1 Pro weights are NOT open. OpenRouter lists S2.1 Pro release as 2026-07-29 (API availability there). Predecessor OpenAudio S1 (`s1`) still supported.","verified":"2026-09-29","body_md":"Sources: https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits , https://huggingface.co/fishaudio/s2-pro , https://fish.audio/blog/s2-1-pro-free-api/","page_url":"https://postcutoff.com/m/fish-audio-s2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":765,"you_url":null,"briefings":null},{"id":"gpt-5-4","name":"GPT-5.4","org":"OpenAI","family":"GPT-5","released":"2026-03-05","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1050000,"max_output":128000,"knowledge_cutoff":"2025-08","pricing":{"input":2.5,"cached_input":0.25,"output":15,"unit_note":"prompts >272K input tokens billed 2x input / 1.5x output","unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2.50 in, $15 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.4","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.4"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.4","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.4","url":"https://openrouter.ai/openai/gpt-5.4"},{"provider":"Web app","url":"https://chatgpt.com"}],"capabilities":[{"name":"Tool search and computer use","detail":"Launched together with API tool search and computer-use support.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/changelog"},{"name":"1M context","detail":"1.05M context window with 128K output; effort defaults to none.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.4"}],"entry":null,"notes":"Snapshot gpt-5.4-2026-03-05. Variants gpt-5.4-pro ($30/$180), gpt-5.4-mini ($0.75/$4.50), gpt-5.4-nano ($0.20/$1.25) are also on the pricing page and OpenRouter.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.4\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.4\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-4/","events_after":846,"major_after":213,"historic_after":42,"missing_at_launch":79,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-5-4/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-08-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-08-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-08-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-08.md"}},{"id":"phi-4-reasoning-vision-15b","name":"Phi-4-Reasoning-Vision-15B","org":"Microsoft","family":"Phi-4","released":"2026-03-04","status":"current","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":16384,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B"},{"provider":"Microsoft Foundry","url":"https://aka.ms/Phi-4-r-v-foundry","docs":"https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-phi-4-reasoning-vision-to-microsoft-foundry/4499154"}],"capabilities":[{"name":"Hybrid think / no-think vision reasoning","detail":"Automatically chooses direct answers for perception tasks and long chain-of-thought only for math/science/diagram problems.","first":false,"discovered":"launch","source":"https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B"},{"name":"GUI grounding for computer-use agents","detail":"Dynamic-resolution SigLIP-2 encoder (up to 3,600 visual tokens) with strengths in GUI grounding for computer-use agents.","first":false,"discovered":"launch","source":"https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B"}],"entry":null,"notes":"Newest Phi model found (Mar 2026). Foundry model id and pricing not verified. Microsoft MAI models (MAI-Image-2/2.5, MAI-Voice-2, MAI-Transcribe-2, MAI-Thinking-1) are in Foundry but not covered by a file here.","verified":"2026-09-29","body_md":"Compact open multimodal reasoning model for visual math/science and screen understanding.\n\n```python\nfrom transformers import AutoProcessor, AutoModelForCausalLM\nmodel = AutoModelForCausalLM.from_pretrained(\"microsoft/Phi-4-reasoning-vision-15B\", trust_remote_code=True, device_map=\"auto\")\n```\n\nSources: https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B","page_url":"https://postcutoff.com/m/phi-4-reasoning-vision-15b/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":767,"you_url":"https://postcutoff.com/you/phi-4-reasoning-vision-15b/","briefings":null},{"id":"songgeneration-2","name":"SongGeneration 2 (LeVo 2)","org":"Tencent AI Lab","family":"SongGeneration / LeVo","released":"2026-03-01","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"custom Tencent terms (GitHub showed NOASSERTION; not verified)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face (v2-large checkpoint, uploader account)","url":"https://huggingface.co/lglg666/SongGeneration-v2-large"},{"provider":"Hugging Face (official org repo; returned 401 on 2026-09-29)","url":"https://huggingface.co/tencent/SongGeneration"}],"capabilities":[{"name":"Hybrid LLM-diffusion full songs up to 4:30","detail":"4B-parameter model generating complete songs up to 4 min 30 s with vocals + accompaniment, instrumental-only, a cappella or dual-track (separated) output; multilingual lyrics (Chinese, English, Spanish, Japanese and more).","first":false,"discovered":"launch","source":"https://github.com/vllm-project/vllm-omni/issues/3390"},{"name":"Hierarchical semantic planning + track-specific refinement","detail":"LeVo 2 paper: semantic planning precedes per-track refinement to keep vocal-instrument coordination while improving acoustics; progressive post-training with automatic quality tiers.","first":false,"discovered":"later","source":"https://arxiv.org/abs/2606.30642"}],"entry":null,"notes":"Released 2026-03-01 (per vLLM-Omni model request citing the official repo). Reported lyric accuracy PER 8.55% vs Suno v5 12.4% and Mureka v8 9.96% (secondary source gaga.art, not verified). As of 2026-09-29 the official GitHub repo github.com/tencent-ailab/SongGeneration returns 404 and the tencent/SongGeneration HF repo returns 401 (apparently removed/made private; community forks and reuploads exist, e.g. Pinokio notes); lglg666/SongGeneration-v2-large (created 2026-02-15, license 'unknown') is still public. Treat availability and license as unverified. Demo: https://levo-demo.github.io/levo_v2_demo/","verified":null,"body_md":"Tencent's open song model family (v1 LeVo, arXiv 2506.07520; v2 LeVo 2, arXiv 2606.30642). Availability is uncertain after the official repos went offline.","page_url":"https://postcutoff.com/m/songgeneration-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":768,"you_url":null,"briefings":null},{"id":"grok-4-20","name":"Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)","org":"xAI","family":"Grok 4","released":"2026-03","status":"legacy","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":1.25,"cached_input":0.2,"output":2.5,"input_over_200k":2.5,"output_over_200k":5,"unit":"per 1M tokens (USD)","source":"https://docs.x.ai/docs/models"},"price_line":"$1.25 in, $2.50 out per 1M tokens","access":[{"provider":"xAI API","model_id":"grok-4.20-0309-reasoning","endpoint":"https://api.x.ai/v1/chat/completions","docs":"https://docs.x.ai/docs/models"},{"provider":"xAI API (non-reasoning)","model_id":"grok-4.20-0309-non-reasoning","endpoint":"https://api.x.ai/v1/chat/completions"},{"provider":"xAI API (multi-agent)","model_id":"grok-4.20-multi-agent-0309","docs":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"provider":"OpenRouter","model_id":"x-ai/grok-4.20","url":"https://openrouter.ai/x-ai/grok-4.20"},{"provider":"OpenRouter (multi-agent)","model_id":"x-ai/grok-4.20-multi-agent","url":"https://openrouter.ai/x-ai/grok-4.20-multi-agent"}],"capabilities":[{"name":"Multi-agent model variant","detail":"Dedicated API id that runs parallel collaborating agents (4 at low/medium effort, 16 at high/xhigh) that search and cross-check before synthesizing an answer.","first":false,"discovered":"launch","source":"https://docs.x.ai/developers/model-capabilities/text/multi-agent"},{"name":"Reasoning and non-reasoning twin ids","detail":"Same snapshot (0309) offered as separate reasoning and non-reasoning model ids.","first":false,"discovered":"launch","source":"https://docs.x.ai/docs/models"}],"entry":null,"notes":"Snapshot ids dated 0309. xAI docs list 1M context; OpenRouter lists 2M. Superseded by Grok 4.5-4.7; logprobs unsupported.","verified":"2026-09-29","body_md":"Previous-generation Grok, notable for its multi-agent variant for deep research.\n\n```bash\ncurl https://api.x.ai/v1/chat/completions -H \"Authorization: Bearer $XAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"grok-4.20-0309-reasoning\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.x.ai/docs/models , https://docs.x.ai/developers/model-capabilities/text/multi-agent","page_url":"https://postcutoff.com/m/grok-4-20/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":747,"you_url":null,"briefings":null},{"id":"gemini-3-1-flash-image","name":"Nano Banana 2 (Gemini 3.1 Flash Image)","org":"Google DeepMind","family":"Gemini Image","released":"2026-02-26","status":"deprecated","type":"image-gen","modality_in":["text","image","video","pdf"],"modality_out":["image","text"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":32768,"knowledge_cutoff":null,"pricing":{"input":0.5,"output":3,"per_image_1k":0.067,"per_image_4k":0.151,"unit":"per 1M tokens text/image input and text output; image output $60 per 1M tokens = $0.045 (0.5K) / $0.067 (1K) / $0.101 (2K) / $0.151 (4K) per image","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.50 in, $3 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.1-flash-image","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-flash-image"},{"provider":"OpenRouter","model_id":"google/gemini-3.1-flash-image"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Pro quality at Flash speed","detail":"Brings Nano Banana Pro world knowledge, reasoning and quality to a fast Flash model.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/"},{"name":"Image search grounding","detail":"Uses real-time web/image search to render real subjects accurately; supports thinking.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image"},{"name":"Text rendering and in-image translation","detail":"Legible text for marketing assets and translation of text inside images.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/"},{"name":"Extreme aspect ratios and 512px-4K","detail":"0.5K/1K/2K/4K outputs and 1:4, 4:1, 1:8, 8:1 ratios; consistency of up to 5 characters and 14 objects.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image"}],"entry":null,"notes":"Deprecated on 2026-10-06 (no shutdown date announced) in favour of gemini-nano-banana-2.1 (Gemini API changelog). Preview id gemini-3.1-flash-image-preview (2026-02-26, still served); stable id GA 2026-05-28. Replacement for gemini-2.5-flash-image and Imagen 4. OpenRouter lists 131k context for stable id.","verified":"2026-10-07","body_md":"Default Google image generation/editing model: fast, text-accurate, search-grounded.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"A poster that says HELLO WORLD in neon letters\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [launch blog](https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/).","page_url":"https://postcutoff.com/m/gemini-3-1-flash-image/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":771,"you_url":null,"briefings":null},{"id":"gpt-audio-1-5","name":"GPT-Audio-1.5 (and gpt-audio / gpt-audio-mini)","org":"OpenAI","family":"GPT Audio","released":"2026-02-23","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":16384,"knowledge_cutoff":"2024-09","pricing":{"text_input":2.5,"text_output":10,"audio_input":32,"audio_output":64,"unit":"per 1M tokens (USD); gpt-audio same; gpt-audio-mini audio $10 / $20, text $0.60 / $2.40","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2.50 text in, $10 text out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-audio-1.5","endpoint":"https://api.openai.com/v1/chat/completions","docs":"https://developers.openai.com/api/docs/models/gpt-audio-1.5"},{"provider":"OpenAI API","model_id":"gpt-audio-mini","endpoint":"https://api.openai.com/v1/chat/completions","docs":"https://developers.openai.com/api/docs/models/gpt-audio-mini"}],"capabilities":[{"name":"Audio in / audio out over Chat Completions","detail":"Non-realtime REST alternative to the Realtime API: send audio and receive spoken audio plus text in one Chat Completions call, with streaming and function calling.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-audio-1.5"}],"entry":null,"notes":"gpt-audio-1.5 released 2026-02-23 with gpt-realtime-1.5. Older gpt-audio (2025) and gpt-audio-mini (2025-10-06) were deprecated 2026-07-20 with shutdown 2027-01-20 (replacement gpt-audio-1.5); gpt-4o-audio-preview was shut down 2026-05-12. Chat Completions only (not Responses).","verified":"2026-09-29","body_md":"Sources:\n- https://developers.openai.com/api/docs/models/gpt-audio-1.5\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations","page_url":"https://postcutoff.com/m/gpt-audio-1-5/","events_after":897,"major_after":238,"historic_after":51,"missing_at_launch":122,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-09.md"}},{"id":"gpt-realtime-1-5","name":"GPT-Realtime-1.5","org":"OpenAI","family":"GPT Realtime","released":"2026-02-23","status":"legacy","type":"audio/speech","modality_in":["text","audio","image"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":32000,"max_output":4096,"knowledge_cutoff":"2024-09","pricing":{"text_input":4,"text_output":16,"audio_input":32,"audio_output":64,"cached_input":0.4,"image_input":5,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$4 text in, $16 text out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-realtime-1.5","endpoint":"wss://api.openai.com/v1/realtime","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-1.5"}],"capabilities":[{"name":"Non-reasoning voice agent model","detail":"Speech-to-speech model for voice agents and customer support with function calling and prompt caching; cheaper text output ($16 vs $24/1M) than the reasoning gpt-realtime-2.x models.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime-1.5"}],"entry":null,"notes":"Released 2026-02-23 alongside gpt-audio-1.5 (Chat Completions). Docs still call it 'our flagship audio model for voice agents', but gpt-realtime-2 (May 2026) and gpt-realtime-2.1 (July 2026) supersede it; not deprecated as of 2026-09-29. It is the named replacement for the gpt-4o-realtime-preview models shut down 2026-05-12.","verified":"2026-09-29","body_md":"Sources:\n- https://developers.openai.com/api/docs/models/gpt-realtime-1.5\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations","page_url":"https://postcutoff.com/m/gpt-realtime-1-5/","events_after":897,"major_after":238,"historic_after":51,"missing_at_launch":122,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-realtime-1-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-09-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-09-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-09-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-09.md"}},{"id":"gemini-3-1-pro","name":"Gemini 3.1 Pro","org":"Google DeepMind","family":"Gemini 3","released":"2026-02-19","status":"preview","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":null,"pricing":{"input":2,"output":12,"unit":"per 1M tokens (Standard, prompts <=200k; >200k: $4.00 in / $18.00 out)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$2 in, $12 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3.1-pro-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-pro"},{"provider":"OpenRouter","model_id":"google/gemini-3.1-pro-preview"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Novel-pattern reasoning","detail":"Verified 77.1% on ARC-AGI-2, more than double Gemini 3 Pro.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/"},{"name":"Custom-tools agent variant","detail":"Separate id gemini-3.1-pro-preview-customtools tuned for agentic workflows using custom tools and bash.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview"},{"name":"Code-generated visuals","detail":"Showcased animated SVG generation, live dashboards and interactive 3D experiences from prompts.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/"}],"entry":null,"notes":"The newest Pro-tier model callable in the Gemini API (preview). Gemini 3.5 Pro, announced at I/O 2026, has not shipped; Google's new top model is Gemini 4 Argon (announced 2026-09-30, which replaces the '3.x Pro' naming), but at launch it was open only to the Fairwind cyber-defense program, with paid API access announced as next and no date. Predecessor gemini-3-pro-preview is shut down. Newer 3.5+ Flash models beat it on many agentic and coding benchmarks at lower cost. Vertex model id not verified.","verified":"2026-09-29","body_md":"Google's Pro-tier reasoning model (preview), the newest one in the Gemini API until Gemini 4 Argon reaches it. Good for hard reasoning, long-context analysis and complex code, though 3.8 Flash is often better value.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Prove that sqrt(2) is irrational\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/), [Gemini 4 Argon announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/).","page_url":"https://postcutoff.com/m/gemini-3-1-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":776,"you_url":"https://postcutoff.com/you/gemini-3-1-pro/","briefings":null},{"id":"lyria-3","name":"Lyria 3 (Clip / Pro)","org":"Google DeepMind","family":"Lyria","released":"2026-02-18","status":"legacy","type":"music","modality_in":["text","image"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_clip":0.04,"per_song":0.08,"unit":"USD per generation: Lyria 3 Clip (30 s clip) $0.04; Lyria 3 Pro (full song, up to ~3 min) $0.08. Same on Gemini API and Vertex.","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.04 per clip","access":[{"provider":"Gemini API (Interactions API)","model_id":"lyria-3-clip-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","docs":"https://ai.google.dev/gemini-api/docs/music-generation"},{"provider":"Gemini API (Interactions API)","model_id":"lyria-3-pro-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"lyria-3-pro-preview","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"lyria-3-clip-preview"},{"provider":"OpenRouter","model_id":"google/lyria-3-pro-preview"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Songs with vocals and auto-written lyrics in the Gemini app","detail":"30-second tracks with vocals and lyrics from a text prompt, photo or video, with Nano Banana cover art; 8 languages (en, de, es, fr, hi, ja, ko, pt); 18+ only.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/products/gemini-app/lyria-3/"},{"name":"SynthID watermark + detection in Gemini","detail":"All outputs carry SynthID; the Gemini app can check whether uploaded audio was generated with Google AI via SynthID.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/products/gemini-app/lyria-3/"},{"name":"Full songs with structure control (Lyria 3 Pro)","detail":"Tracks up to ~3 minutes (184 s max on Vertex) with control over intros, verses, choruses and bridges, duration, BPM and intensity; 44.1 kHz, 192 kbps MP3; C2PA content credentials and vocal-likeness filtering.","first":false,"discovered":"launch","source":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3"}],"entry":"2026-02-18-google-lyria-3-gemini-app","notes":"Lyria 3 launched 2026-02-18 in the Gemini app (30 s clips) and YouTube Dream Track; Lyria 3 Pro and the developer previews (lyria-3-clip-preview, lyria-3-pro-preview) followed on 2026-03-25 (Gemini API, AI Studio, Vertex public preview, Google Vids, ProducerAI). Superseded by Lyria 3.5 (lyria-3.5, GA 2026-09-03); Gemini API pricing page now lists both as 'Lyria 3 legacy models'; no shutdown date announced. Artist names in prompts are treated as broad inspiration only.","verified":"2026-09-29","body_md":"Google's first Lyria generation with vocals and lyrics; now superseded by [Lyria 3.5](lyria-3-5.md) but still served as previews.\n\nSources: [Lyria 3 in Gemini](https://blog.google/innovation-and-ai/products/gemini-app/lyria-3/), [Lyria 3 Pro](https://blog.google/innovation-and-ai/technology/ai/lyria-3-pro/), [Vertex model page](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-3), [Gemini API pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/lyria-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":777,"you_url":null,"briefings":null},{"id":"claude-sonnet-4-6","name":"Claude Sonnet 4.6","org":"Anthropic","family":"Claude 4","released":"2026-02-17","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2025-08","pricing":{"input":3,"output":15,"cache_read":0.3,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$3 in, $15 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-sonnet-4-6","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/sonnet-4-6/overview"},{"provider":"AWS Bedrock (InvokeModel)","model_id":"anthropic.claude-sonnet-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy"},{"provider":"Google Cloud Vertex AI","model_id":"claude-sonnet-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-sonnet-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-sonnet-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-sonnet-4.6","url":"https://openrouter.ai/anthropic/claude-sonnet-4.6"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Human-level computer use on common tasks","detail":"Anthropic cites human-level performance on tasks such as navigating complex spreadsheets and multi-step web forms (OSWorld).","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-6"},{"name":"Beats previous Opus in user preference","detail":"Users preferred it to Opus 4.5 59% of the time on coding, citing less overengineering.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-6"},{"name":"1M context for Sonnet 4.6","detail":"1M-token context window (beta at launch).","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-6"}],"entry":null,"notes":"Last model on the older tokenizer. Adaptive thinking (budget_tokens deprecated). Training data cutoff Jan 2026. Bedrock via InvokeModel. Retirement not sooner than 2027-02-17.","verified":"2026-09-29","body_md":"Legacy Sonnet; migrate to `claude-sonnet-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-sonnet-4-6\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/sonnet-4-6/overview\n- Announcement: https://www.anthropic.com/news/claude-sonnet-4-6\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-sonnet-4-6/","events_after":846,"major_after":213,"historic_after":42,"missing_at_launch":67,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-sonnet-4-6/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-08-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-08-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-08-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-08.md"}},{"id":"xiaomi-robotics-0","name":"Xiaomi-Robotics-0 (4.7B VLA)","org":"Xiaomi","family":"Xiaomi-Robotics","released":"2026-02-12","status":"legacy","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"XiaomiRobotics/Xiaomi-Robotics-0-Pretrain","url":"https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-0-Pretrain"},{"provider":"Project page","url":"https://xiaomi-robotics-0.github.io"}],"capabilities":[{"name":"Real-time asynchronous execution on a consumer GPU","detail":"Post-trained for asynchronous execution with aligned timesteps between consecutive action chunks, so rollouts stay smooth despite inference latency; runs on a consumer-grade GPU (per paper).","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2602.12684"},{"name":"Strong open sim-benchmark results","detail":"LIBERO 98.7% avg; SimplerEnv Visual Matching 85.5%, Visual Aggregation 74.7%, WidowX 79.2%; CALVIN avg length 4.75 (ABC-D) / 4.80 (ABCD-D) (authors).","first":false,"discovered":"launch","source":"https://xiaomi-robotics-0.github.io"}],"entry":null,"notes":"4.7B parameters, Qwen3-VL-4B-Instruct backbone; pretrained on cross-embodiment robot trajectories plus vision-language data. Real-robot evals: Lego disassembly and towel folding (bimanual). Checkpoints: -Pretrain, -LIBERO, -Calvin-ABC_D, -Calvin-ABCD_D, -SimplerEnv-WidowX, -SimplerEnv-Google-Robot (HF, 2026-02-10). Paper arXiv 2602.12684 (2026-02-13). Superseded by xiaomi-robotics-1 (July 2026).","verified":"2026-09-29","body_md":"Xiaomi's first open VLA. Use [xiaomi-robotics-1](xiaomi-robotics-1.md) for new work.\n\nSources: [project page](https://xiaomi-robotics-0.github.io), [arXiv 2602.12684](https://arxiv.org/abs/2602.12684), [HF](https://huggingface.co/XiaomiRobotics).","page_url":"https://postcutoff.com/m/xiaomi-robotics-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":781,"you_url":null,"briefings":null},{"id":"claude-opus-4-6","name":"Claude Opus 4.6","org":"Anthropic","family":"Claude 4","released":"2026-02-05","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":128000,"knowledge_cutoff":"2025-05","pricing":{"input":5,"output":25,"cache_read":0.5,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-4-6","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-4-6/overview"},{"provider":"AWS Bedrock (InvokeModel)","model_id":"anthropic.claude-opus-4-6-v1","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-4-6","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-4.6","url":"https://openrouter.ai/anthropic/claude-opus-4.6"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"1M-token context for Opus","detail":"First Opus with a 1M-token context window (launched in beta). Scored 76% on MRCR v2 long-context retrieval versus 18.5% for Sonnet 4.5.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-6"},{"name":"Adaptive thinking","detail":"Introduced adaptive thinking: the model decides when and how much to think, steered by effort.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-6"},{"name":"Agent teams","detail":"Research preview of multiple Claude instances coordinating in parallel (in Claude Code).","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-6"},{"name":"Knowledge work (GDPval-AA)","detail":"About 144 Elo above GPT-5.2 on GDPval-AA. Also led Terminal-Bench 2.0 at launch.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-6"}],"entry":null,"notes":"First dateless-ID Opus; adaptive thinking (budget_tokens deprecated). Training data cutoff Aug 2025. Bedrock via InvokeModel only. Retirement not sooner than 2027-02-05.","verified":"2026-09-29","body_md":"Legacy Opus; migrate to `claude-opus-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-4-6\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-4-6/overview\n- Announcement: https://www.anthropic.com/news/claude-opus-4-6\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-4-6/","events_after":863,"major_after":219,"historic_after":44,"missing_at_launch":73,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-4-6/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-05-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-05-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-05-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-05.md"}},{"id":"gpt-5-3-codex","name":"GPT-5.3-Codex","org":"OpenAI","family":"GPT-5 Codex","released":"2026-02-05","status":"current","type":"code","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":400000,"max_output":128000,"knowledge_cutoff":"2025-08","pricing":{"input":1.75,"cached_input":0.175,"output":14,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$1.75 in, $14 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-5.3-codex","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-5.3-codex"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-5.3-codex","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-5.3-codex","url":"https://openrouter.ai/openai/gpt-5.3-codex"},{"provider":"Codex (ChatGPT)","url":"https://chatgpt.com/codex"}],"capabilities":[{"name":"Agentic coding specialist","detail":"Codex-tuned GPT-5.3 for long-running software engineering (Codex app/CLI/IDE and API).","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.3-codex"},{"name":"Responses-only","detail":"Available only through the Responses API; effort low/medium/high/xhigh.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-5.3-codex"}],"entry":null,"notes":"Latest codex-specific API id on the pricing page. Released in Codex Feb 5 2026; API access followed later (Azure version 2026-02-24). GPT-6 Sol is now positioned for coding.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-5.3-codex\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-5.3-codex\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-5-3-codex/","events_after":846,"major_after":213,"historic_after":42,"missing_at_launch":56,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-08-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-08-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-08-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-08.md"}},{"id":"rime-arcana-v3","name":"Rime Arcana v3 / v3 Turbo","org":"Rime","family":"Arcana","released":"2026-02-04","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Rime API","model_id":"arcana","endpoint":"https://users.rime.ai (geo endpoints users-west.rime.ai, users-east.rime.ai; HTTP and WebSocket JSON streaming)","docs":"https://docs.rime.ai/"},{"provider":"Together AI","model_id":"Rime Arcana V3 / Arcana V3 Turbo (dedicated endpoints)","url":"https://www.together.ai/models/rime-arcana-v3-turbo"},{"provider":"Telnyx","url":"https://telnyx.com/release-notes/rime-arcana-v3-voices"},{"provider":"On-prem","url":"https://www.rime.ai/resources/arcana-v3"}],"capabilities":[{"name":"Native code-switching across 10 languages","detail":"One voice switches mid-conversation among English, Hindi, Spanish, Arabic, French, Portuguese, German, Japanese, Hebrew and Tamil (Together AI lists 11 languages); word-level timestamps.","first":false,"discovered":"launch","source":"https://www.rime.ai/resources/arcana-v3"},{"name":"Enterprise latency and on-prem scale","detail":"~120 ms on-prem model latency, ~200 ms TTFB via cloud API, 100+ concurrent generations per machine; Rapidata listener tests (US) preferred it 61-64% of the time over ElevenLabs Turbo v2.5, Google Chirp and Cartesia Sonic (vendor-run).","first":false,"discovered":"launch","source":"https://www.rime.ai/resources/arcana-v3"}],"entry":null,"notes":"Calling the existing `arcana` model id automatically serves v3. Arcana V3 Turbo is the low-latency variant (Together AI: ~120 ms time-to-first-audio, $10 per 1M characters plus GPU-hour on dedicated endpoints). Earlier: Arcana (Apr 2025), Arcana v2. Rime's own per-character price not verified.","verified":"2026-09-29","body_md":"Rime's flagship TTS for enterprise voice agents (call centres, scheduling).\n\nLaunch video: [rime-arcana-v3-launch](../videos/rime-arcana-v3-launch.md).\n\nSources: https://www.rime.ai/resources/arcana-v3 , https://x.com/rimelabs/status/2019099676939813306 , https://www.together.ai/blog/rime-arcana-v3-turbo-and-rime-arcana-v3-now-available-on-together-ai","page_url":"https://postcutoff.com/m/rime-arcana-v3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":790,"you_url":null,"briefings":null},{"id":"voxtral-transcribe-2","name":"Voxtral Transcribe 2 (Mini Transcribe V2 + Voxtral Realtime)","org":"Mistral AI","family":"Voxtral","released":"2026-02-04","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute_batch":0.003,"per_minute_realtime":0.006,"unit":"USD per minute of audio (voxtral-mini-2602 batch / voxtral-mini-transcribe-realtime-2602)","source":"https://mistral.ai/news/voxtral-transcribe-2"},"price_line":"$0.003 per minute, batch","access":[{"provider":"Mistral API (batch)","model_id":"voxtral-mini-2602","endpoint":"https://api.mistral.ai/v1/audio/transcriptions","docs":"https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-26-02"},{"provider":"Mistral API (realtime)","model_id":"voxtral-mini-transcribe-realtime-2602","docs":"https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-realtime-26-02"},{"provider":"Hugging Face (Realtime, open weights)","model_id":"mistralai/Voxtral-Mini-4B-Realtime-2602","url":"https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"Open-weight realtime ASR under 200 ms","detail":"Voxtral Realtime (4B, Apache 2.0) reaches sub-200 ms latency; at 480 ms delay Mistral reports 1-2% WER.","first":false,"discovered":"launch","source":"https://mistral.ai/news/voxtral-transcribe-2"},{"name":"Cheap batch transcription with diarization","detail":"Mini Transcribe V2: ~4% WER on FLEURS at $0.003/min with speaker diarization, word timestamps, context biasing (up to 100 terms) and audio up to 3 hours.","first":false,"discovered":"launch","source":"https://mistral.ai/news/voxtral-transcribe-2"}],"entry":"2026-03-23-mistral-voxtral-tts","notes":"13 languages (en, zh, hi, es, ar, fr, pt, ru, de, ja, ko, it, nl). Batch model is API-only ('Premier' license); Realtime has open weights. Replaced voxtral-mini-2507 / Voxtral Mini Transcribe (deprecated 2026-02-27, retired 2026-05-31). Tech report arXiv 2602.11298. Accuracy claims are Mistral's.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.mistral.ai/v1/audio/transcriptions -H \"Authorization: Bearer $MISTRAL_API_KEY\" \\\n  -F model=voxtral-mini-2602 -F file=@audio.mp3\n```\n\nSources: https://mistral.ai/news/voxtral-transcribe-2 · https://docs.mistral.ai/models/overview","page_url":"https://postcutoff.com/m/voxtral-transcribe-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":790,"you_url":null,"briefings":null},{"id":"kling-3-0","name":"Kling 3.0 (VIDEO 3.0 / 3.0 Omni)","org":"Kuaishou","family":"Kling 3","released":"2026-02","status":"current","type":"video-gen","modality_in":["text","image","video"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Kling AI API","endpoint":"https://api-singapore.klingai.com","docs":"https://kling.ai/document-api/quickStart/productIntroduction/overview"},{"provider":"fal.ai","model_id":"fal-ai/kling-video/v3/standard/text-to-video","docs":"https://fal.ai/models/fal-ai/kling-video/v3/standard/text-to-video/api"},{"provider":"Web app","url":"https://kling.ai"}],"capabilities":[{"name":"Multi-shot storyboards","detail":"Generates multi-shot narrative sequences in one job, with storyboard control over shots.","first":false,"discovered":"launch","source":"https://kling.ai/quickstart/klingai-video-3-model-user-guide"},{"name":"Native multilingual audio","detail":"Native audio (dialogue/SFX) generated with the video, multilingual; clips up to 15 s.","first":false,"discovered":"launch","source":"https://kling.ai/quickstart/klingai-video-3-model-user-guide"},{"name":"Unified Omni model with element consistency","detail":"VIDEO 3.0 Omni (successor of O1) unifies generation and editing with stronger element/character consistency; IMAGE 3.0 / 3.0 Omni siblings.","first":false,"discovered":"launch","source":"https://kling.ai/quickstart/klingai-video-3-model-user-guide"}],"entry":null,"notes":"Released early Feb 2026 (official guide says Feb 6; other sources Feb 7). Official API model_name strings not verified (docs are JS-rendered); fal ids verified: kling-video/v3/{standard,pro}/{text,image}-to-video, plus turbo/4K variants. Pricing not verified.","verified":null,"body_md":"Kuaishou's flagship video model with native audio and multi-shot generation. Official API is async (submit task, poll task id).\n\n```bash\n# via fal.ai\ncurl -X POST https://fal.run/fal-ai/kling-video/v3/standard/text-to-video -H \"Authorization: Key $FAL_KEY\" \\\n  -H \"Content-Type: application/json\" -d '{\"prompt\":\"a samurai walks through falling snow, cinematic\",\"duration\":\"5\"}'\n```\n\nSources: https://kling.ai/quickstart/klingai-video-3-model-user-guide , https://kling.ai/document-api/quickStart/productIntroduction/overview , https://fal.ai/models/fal-ai/kling-video/v3/standard/text-to-video/api","page_url":"https://postcutoff.com/m/kling-3-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":769,"you_url":null,"briefings":null},{"id":"qwen3-asr","name":"Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner","org":"Alibaba (Qwen)","family":"Qwen3-ASR","released":"2026-01-29","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"Qwen/Qwen3-ASR-1.7B","url":"https://huggingface.co/Qwen/Qwen3-ASR-1.7B"},{"provider":"Hugging Face","model_id":"Qwen/Qwen3-ASR-0.6B","url":"https://huggingface.co/Qwen/Qwen3-ASR-0.6B"},{"provider":"Hugging Face","model_id":"Qwen/Qwen3-ForcedAligner-0.6B","url":"https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B"},{"provider":"GitHub","url":"https://github.com/QwenLM/Qwen3-ASR"}],"capabilities":[{"name":"52 languages/dialects incl. singing and music","detail":"Language ID + ASR for 30 languages and 22 Chinese dialects, robust on songs/music; built on Qwen3-Omni audio understanding; vLLM batch and streaming inference, timestamp prediction via ForcedAligner.","first":false,"discovered":"launch","source":"https://github.com/QwenLM/Qwen3-ASR"},{"name":"Beats Whisper-large-v3 on Chinese","detail":"Self-reported WER e.g. AISHELL-2 2.71 vs 5.06 (Whisper-large-v3); Cantonese CV-yue 7.57 vs 11.36 (GPT-4o-Transcribe).","first":false,"discovered":"launch","source":"https://github.com/QwenLM/Qwen3-ASR"}],"entry":"2026-01-22-qwen3-tts-asr-open-weights","notes":"Native Transformers (-hf repos) support added 2026-06-26. Hosted ASR is now Qwen-Audio-3.x-ASR (see qwen-audio-3-1-asr).","verified":"2026-09-29","body_md":"Apache-2.0 speech recognition models for self-hosting.\n\nSources: https://github.com/QwenLM/Qwen3-ASR","page_url":"https://postcutoff.com/m/qwen3-asr/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":793,"you_url":null,"briefings":null},{"id":"ace-step-1-5","name":"ACE-Step 1.5 (incl. 1.5 XL)","org":"ACE Studio & StepFun","family":"ACE-Step","released":"2026-01-28","status":"current","type":"music","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/ACE-Step/Ace-Step1.5"},{"provider":"Hugging Face (XL 4B DiT)","url":"https://huggingface.co/ACE-Step/acestep-v15-xl-sft"},{"provider":"GitHub","url":"https://github.com/ace-step/ACE-Step-1.5"},{"provider":"Web app","url":"https://acemusic.ai"}],"capabilities":[{"name":"Full songs in seconds on consumer hardware","detail":"10 s to 10 min of music; under 2 s per song on an A100 and under 10 s on an RTX 3090; standard models run in <4 GB VRAM with offload (XL: >=12 GB, 20 GB recommended).","first":false,"discovered":"launch","source":"https://github.com/ace-step/ACE-Step-1.5"},{"name":"LM planner + DiT synthesizer","detail":"A language model (0.6B/1.7B/4B '5Hz LM') turns prompts into a song blueprint that a Diffusion Transformer renders; aligned with 'intrinsic' RL without external reward models.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2602.00744"},{"name":"Editing and personalization toolkit","detail":"Cover generation, repaint/editing, vocal-to-BGM, track separation, multi-track generation, BPM/key extraction and LoRA fine-tuning from ~8 songs (about 1 h on a 12 GB RTX 3090); lyrics in 50+ languages.","first":false,"discovered":"launch","source":"https://github.com/ace-step/ACE-Step-1.5"}],"entry":"2026-01-28-ace-step-1-5","notes":"Checkpoints: acestep-v15-base / -sft / -turbo (plus turbo-shift variants) and, from 2026-04-02, XL (4B DiT) xl-base / xl-sft / xl-turbo; diffusers versions added Apr-Jun 2026. Release date 2026-01-28 is from secondary sources (HF repos created 2026-01-23, arXiv 2602.00744 submitted 2026-01-31). Authors claim quality beyond most commercial models (SongEval above Suno v5 per secondary coverage; not independently verified). Supports Mac, AMD, Intel and CUDA.","verified":"2026-09-29","body_md":"The go-to MIT-licensed local song generator in 2026. Quick start: clone https://github.com/ace-step/ACE-Step-1.5 and follow the README (Gradio UI and API server included).\n\nSources: [GitHub](https://github.com/ace-step/ACE-Step-1.5), [HF](https://huggingface.co/ACE-Step/Ace-Step1.5), [tech report](https://arxiv.org/abs/2602.00744), [project page](https://ace-step.github.io/ace-step-v1.5.github.io/).","page_url":"https://postcutoff.com/m/ace-step-1-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":795,"you_url":null,"briefings":null},{"id":"figure-helix-02","name":"Helix 02","org":"Figure AI","family":"Helix","released":"2026-01-27","status":"legacy","type":"robotics","modality_in":["text","image","video"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"None (runs only on Figure 03 robots; no public API or weights)","url":"https://www.figure.ai/news/helix-02"}],"capabilities":[{"name":"Pixels-to-whole-body control over long horizons","detail":"One visuomotor network links every sensor (vision, touch, proprioception) to every actuator; unloaded and reloaded a dishwasher across a full kitchen in a 4-minute run with walking, manipulation and balance, no resets.","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix-02"},{"name":"System 0 learned whole-body controller","detail":"New 10M-parameter S0 at 1 kHz trained on 1,000+ hours of retargeted human motion and 200,000+ parallel simulated environments, under S1 (200 Hz) and S2 (semantic reasoning).","first":false,"discovered":"launch","source":"https://www.figure.ai/news/helix-02"},{"name":"Tactile and palm-camera policies","detail":"First Figure policies that depend on Figure 03's palm cameras and fingertip tactile sensing for occluded, delicate manipulation.","first":false,"discovered":"launch","source":"https://www.figure.ai/news/helix-02"}],"entry":"2026-01-27-figure-helix-02","notes":"Figure: 'first demonstration of such long horizon, end-to-end pixels-to-whole body control on a humanoid robot' (company claim). Superseded by Helix 2.5 (2026-09-17). No external access.","verified":null,"body_md":"Sources: [Figure: Introducing Helix 02](https://www.figure.ai/news/helix-02), [video](https://www.youtube.com/watch?v=lQsvTrRTBRs).","page_url":"https://postcutoff.com/m/figure-helix-02/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":796,"you_url":null,"briefings":null},{"id":"minimax-speech-2-8","name":"MiniMax Speech 2.8 (HD / Turbo)","org":"MiniMax","family":"MiniMax Speech","released":"2026-01-23","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_million_characters_hd":100,"per_million_characters_turbo":60,"unit":"USD per 1M characters (sync and async T2A). Voice clone 1.5 USD/voice, voice design 3 USD/voice","source":"https://platform.minimax.io/docs/guides/pricing-paygo"},"price_line":"$100 per 1M characters, HD","access":[{"provider":"MiniMax API (T2A HTTP / WebSocket / async)","model_id":"speech-2.8-hd","endpoint":"https://api.minimax.io","docs":"https://platform.minimax.io/docs/api-reference/speech-t2a-http"},{"provider":"MiniMax API","model_id":"speech-2.8-turbo","endpoint":"https://api.minimax.io","docs":"https://platform.minimax.io/docs/api-reference/speech-t2a-http"},{"provider":"Web app (MiniMax Audio)","url":"https://www.minimax.io/audio"}],"capabilities":[{"name":"Sound tags","detail":"Natural sound tags (non-verbal cues) in ultra-realistic HD speech.","first":false,"discovered":"launch","source":"https://platform.minimax.io/docs/release-notes/models"},{"name":"40 languages, 7 emotions","detail":"40 languages plus specified dialects, 7 emotions; rapid voice cloning and text-described voice design.","first":false,"discovered":"launch","source":"https://platform.minimax.io/docs/guides/models-intro"},{"name":"Streaming and long-form modes","detail":"Sync HTTP, WebSocket and bidirectional streaming (pipe LLM tokens straight to speech), plus async jobs up to 1M characters.","first":false,"discovered":"launch","source":"https://platform.minimax.io/docs/guides/pricing-paygo"}],"entry":null,"notes":"speech-2.6 and speech-02 are legacy at the same prices. MiniMax also offers ASR (0.38 USD/hour). Music: music-3.0 API closed to new users from 2026-08-20; open weights MiniMax-Music3 on HF.","verified":"2026-09-29","body_md":"MiniMax's current TTS: expressive multilingual voices, cloning and real-time streaming for agents, audiobooks and dubbing.\n\nSources: https://platform.minimax.io/docs/guides/models-intro · https://platform.minimax.io/docs/guides/pricing-paygo · https://platform.minimax.io/docs/api-reference/speech-t2a-http","page_url":"https://postcutoff.com/m/minimax-speech-2-8/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":798,"you_url":null,"briefings":null},{"id":"qwen3-tts","name":"Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)","org":"Alibaba (Qwen)","family":"Qwen3-TTS","released":"2026-01-22","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice","url":"https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice"},{"provider":"Hugging Face","model_id":"Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign","url":"https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign"},{"provider":"GitHub","url":"https://github.com/QwenLM/Qwen3-TTS"},{"provider":"Alibaba Cloud Model Studio","model_id":"qwen3-tts-flash","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen-tts"},{"provider":"Alibaba Cloud Model Studio (instruct / voice design / voice clone)","model_id":"qwen3-tts-instruct-flash","docs":"https://www.alibabacloud.com/help/en/model-studio/qwen-tts"}],"capabilities":[{"name":"Open-weights voice design and 3-second cloning","detail":"Voice design from natural-language descriptions and voice cloning from ~3 s of audio, in 10 languages (zh, en, ja, ko, de, fr, ru, pt, es, it).","first":false,"discovered":"launch","source":"https://github.com/QwenLM/Qwen3-TTS"},{"name":"97 ms streaming latency","detail":"12 Hz multi-codebook tokenizer; first audio packet after a single input character, end-to-end latency as low as 97 ms; one model for streaming and non-streaming.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2601.15621"}],"entry":"2026-01-22-qwen3-tts-asr-open-weights","notes":"HF repos: Qwen3-TTS-12Hz-{1.7B,0.6B}-{Base,CustomVoice}, 1.7B-VoiceDesign, Qwen3-TTS-Tokenizer-12Hz. API snapshots: qwen3-tts-flash (=2025-11-27), qwen3-tts-flash-2025-09-18, qwen3-tts-instruct-flash-2026-01-26, qwen3-tts-vd-2026-01-26 (voice design), qwen3-tts-vc-2026-01-22 (voice clone). Superseded in Alibaba's hosted lineup by Qwen-Audio-3.0-TTS (Jul 2026) and Qwen-Audio-3.1-TTS (Sep 2026). API pricing not verified.","verified":"2026-09-29","body_md":"Apache-2.0 multilingual TTS you can run locally; hosted versions on Model Studio.\n\nSources: https://github.com/QwenLM/Qwen3-TTS · https://arxiv.org/abs/2601.15621 · https://www.alibabacloud.com/help/en/model-studio/qwen-tts","page_url":"https://postcutoff.com/m/qwen3-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":798,"you_url":null,"briefings":null},{"id":"flux-2-klein","name":"FLUX.2 [klein] (4B / 9B)","org":"Black Forest Labs","family":"FLUX.2","released":"2026-01-14","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":true,"model_license":"apache-2.0 (4B); FLUX Non-Commercial License (9B)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.014,"unit":"per image from (4B; 9B from 0.015; text-to-image and editing)","source":"https://docs.bfl.ai/quick_start/pricing"},"price_line":"$0.014 per image","access":[{"provider":"BFL API","model_id":"flux-2-klein-4b","endpoint":"https://api.bfl.ai/v1/flux-2-klein-4b","docs":"https://docs.bfl.ai/flux_2/flux2_overview"},{"provider":"BFL API","model_id":"flux-2-klein-9b","endpoint":"https://api.bfl.ai/v1/flux-2-klein-9b","docs":"https://docs.bfl.ai/flux_2/flux2_overview"},{"provider":"Hugging Face","url":"https://huggingface.co/black-forest-labs/FLUX.2-klein-4B"},{"provider":"Hugging Face","url":"https://huggingface.co/black-forest-labs/FLUX.2-klein-9B"}],"capabilities":[{"name":"Sub-second generation and editing","detail":"Size-distilled FLUX.2 variants aimed at sub-second inference for both text-to-image and editing.","first":false,"discovered":"launch","source":"https://docs.bfl.ai/flux_2/flux2_overview"},{"name":"Apache-2.0 open weights (4B)","detail":"4B checkpoint is Apache 2.0 - commercially usable open weights; base (undistilled) checkpoints published for fine-tuning/LoRA training.","first":false,"discovered":"launch","source":"https://huggingface.co/black-forest-labs/FLUX.2-klein-4B"},{"name":"KV-cached 9B variant","detail":"flux-2-klein-9b-preview / FLUX.2-klein-9b-kv (Mar 2026) add KV caching for faster multi-reference editing.","first":false,"discovered":"later","source":"https://docs.bfl.ai/flux_2/flux2_overview"}],"entry":null,"notes":"Snapshots flux-2-klein-9b (fixed) and flux-2-klein-9b-preview (latest, KV caching). HF also hosts -base, fp8 and nvfp4 variants. Release date = HF repo creation date.","verified":"2026-09-29","body_md":"Small, fast FLUX.2 models for real-time/local use; 4B is Apache-2.0.\n\n```bash\ncurl -X POST https://api.bfl.ai/v1/flux-2-klein-4b -H \"x-key: $BFL_API_KEY\" -H \"Content-Type: application/json\" -d '{\"prompt\":\"minimal poster, the word KLEIN in bold red\"}'\n```\nLocal: `diffusers` with `black-forest-labs/FLUX.2-klein-4B`.\n\nSources: https://docs.bfl.ai/flux_2/flux2_overview , https://docs.bfl.ai/quick_start/pricing , https://huggingface.co/black-forest-labs","page_url":"https://postcutoff.com/m/flux-2-klein/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":800,"you_url":null,"briefings":null},{"id":"kyutai-pocket-tts","name":"Kyutai Pocket TTS","org":"Kyutai","family":"Pocket TTS","released":"2026-01-13","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"cc-by-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"kyutai/pocket-tts","url":"https://huggingface.co/kyutai/pocket-tts"},{"provider":"Hugging Face (no cloning variant)","model_id":"kyutai/pocket-tts-without-voice-cloning","url":"https://huggingface.co/kyutai"},{"provider":"GitHub / pip","model_id":"pocket-tts","url":"https://github.com/kyutai-labs/pocket-tts"}],"capabilities":[{"name":"100M-param TTS with cloning, real time on CPU","detail":"~200 ms to first audio and ~6x real time on a MacBook Air M4 CPU; streaming, unbounded text length; voice cloning from audio.","first":false,"discovered":"launch","source":"https://huggingface.co/kyutai/pocket-tts"},{"name":"Six languages","detail":"English, French, German, Spanish, Portuguese, Italian (multilingual since 2026-05-04).","first":false,"discovered":"later","source":"https://kyutai.org/blog/"}],"entry":null,"notes":"Gated on HF (accept prohibited-use terms). Training code released 2026-08-25; 2026-09-28 post describes a 'drifting' objective replacing flow matching for the sampler head. `pip install pocket-tts`. Community WebAssembly ports run in-browser.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/kyutai/pocket-tts , https://github.com/kyutai-labs/pocket-tts , https://kyutai.org/blog/","page_url":"https://postcutoff.com/m/kyutai-pocket-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":801,"you_url":null,"briefings":null},{"id":"1x-world-model","name":"1X World Model (1XWM)","org":"1X Technologies","family":"1X World Model","released":"2026-01-12","status":"preview","type":"world-model","modality_in":["text","image","video"],"modality_out":["video","action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Not available (internal; runs NEO policies)","url":"https://www.1x.tech/discover/world-model-self-learning","docs":"https://www.1x.tech/1x-world-model.pdf"}],"capabilities":[{"name":"Video world model used as the robot policy","detail":"Given a text prompt, a 14B generative video model fine-tuned on NEO imagines ~5 s of future video; an inverse-dynamics model converts it into actions executed on NEO (≈11 s per rollout on multi-GPU inference).","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/world-model-self-learning"},{"name":"Learns from human egocentric video","detail":"Trained with ~900 h of egocentric human video plus ~70 h of NEO data (and 400 h of unfiltered robot data for the IDM); generalizes to some objects and motions absent from NEO task data. Grasping ~80% success; pouring 0%; best-of-8 generations raised 'pull tissue' from 30% to 45%.","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/world-model-self-learning"},{"name":"World model for policy evaluation","detail":"The June 2025 version was an action-conditioned simulator used to rank policies without physical tests (1X: 70% world-model accuracy picks the better policy ~90% of the time).","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/redwood-ai-world-model"}],"entry":"2026-01-12-1x-world-model-policy","notes":"Two stages: 1XWM as a policy evaluator (2025-06-16) and as a NEO policy (2026-01-12). No API or weights. TechCrunch coverage: https://techcrunch.com/2026/01/13/neo-humanoid-maker-1x-releases-world-model-to-help-bots-learn-what-they-see/","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/1x-world-model/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":802,"you_url":null,"briefings":null},{"id":"elevenlabs-scribe-v2","name":"Scribe v2 / Scribe v2 Medical","org":"ElevenLabs","family":"Scribe","released":"2026-01-09","status":"current","type":"audio/speech","modality_in":["audio","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.22,"unit":"per hour of audio (Scribe v2 and v2 Medical, API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.22 per hour","access":[{"provider":"ElevenLabs API","model_id":"scribe_v2","endpoint":"https://api.elevenlabs.io/v1/speech-to-text","docs":"https://elevenlabs.io/docs/overview/capabilities/speech-to-text"},{"provider":"ElevenLabs API","model_id":"scribe_v2_medical","endpoint":"https://api.elevenlabs.io/v1/speech-to-text","docs":"https://elevenlabs.io/docs/models"}],"capabilities":[{"name":"Entity detection with timestamps","detail":"Native detection of PII, health and payment entities (56 categories at launch, 65 types per current docs) with exact timestamps.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/introducing-scribe-v2"},{"name":"Keyterm prompting, 32-speaker diarization","detail":"Keyterm prompting (100 terms at launch, now up to 1,000), speaker diarization up to 32 speakers, word timestamps, dynamic audio-event tagging, multi-language audio in one file; 90+ languages.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"},{"name":"Clinical variant","detail":"scribe_v2_medical fine-tuned for clinical audio, HIPAA with BAA; generally available 2026-09-14.","first":false,"discovered":"later","source":"https://elevenlabs.io/docs/changelog"}],"entry":null,"notes":"Launched 2026-01-09; ElevenLabs claims 'the lowest word error rate recorded on industry-standard benchmarks' (FLEURS chart; company claim). Realtime variant in its own file. scribe_v1 (launched 2025-02-26, $0.40/h at launch) is deprecated ('outclassed by v2').","verified":"2026-09-29","body_md":"```bash\ncurl -X POST https://api.elevenlabs.io/v1/speech-to-text -H \"xi-api-key: $ELEVENLABS_API_KEY\" \\\n  -F model_id=scribe_v2 -F file=@audio.mp3\n```\n\nSources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/blog/introducing-scribe-v2 , https://elevenlabs.io/blog/scribe-v2-medical-is-now-available-to-everyone","page_url":"https://postcutoff.com/m/elevenlabs-scribe-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":804,"you_url":null,"briefings":null},{"id":"unifolm-vla-0","name":"UnifoLM-VLA-0 (and UnifoLM-WMA-0)","org":"Unitree Robotics","family":"UnifoLM","released":"2026-01","status":"legacy","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"cc-by-nc-sa-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"unitreerobotics/UnifoLM-VLA-Base","url":"https://huggingface.co/collections/unitreerobotics/unifolm-vla-0"},{"provider":"Hugging Face (world-model-action)","model_id":"unitreerobotics/UnifoLM-WMA-0-Base","url":"https://huggingface.co/unitreerobotics/UnifoLM-WMA-0-Base"},{"provider":"GitHub","url":"https://github.com/unitreerobotics/unifolm-vla"}],"capabilities":[{"name":"Open VLA for general-purpose humanoid manipulation","detail":"Continued pretraining of a VLM (UnifoLM-VLM-Base, Qwen2.5-VL based) on robot manipulation data to turn it into an 'embodied brain'; variants fine-tuned on Unitree open datasets and LIBERO.","first":false,"discovered":"launch","source":"https://huggingface.co/collections/unitreerobotics/unifolm-vla-0"},{"name":"World-model-action architecture (WMA-0)","detail":"UnifoLM-WMA-0 (Sept 2025, Apache-2.0) pairs a world model that predicts future interactions (usable as a simulator) with action generation; Base and Dual variants on HF.","first":false,"discovered":"launch","source":"https://huggingface.co/unitreerobotics/UnifoLM-WMA-0-Base"}],"entry":null,"notes":"HF repos for UnifoLM-VLA-Base created 2026-01-28 (exact announcement day not verified). VLA-0 license CC BY-NC-SA 4.0 (non-commercial); WMA-0 Apache-2.0. Superseded by UnifoLM-WLA-1.0 (2026-09). Unitree also publishes ~200 G1 teleoperation datasets under huggingface.co/unitreerobotics.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/unifolm-vla-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":793,"you_url":null,"briefings":null},{"id":"cosmos-reason-2","name":"Cosmos Reason 2","org":"NVIDIA","family":"Cosmos","released":"2025-12-19","status":"current","type":"multimodal","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"NVIDIA Open Model License (commercial use allowed; HF repo gated - accept terms)","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/Cosmos-Reason2-8B","url":"https://huggingface.co/nvidia/Cosmos-Reason2-8B"},{"provider":"Hugging Face","model_id":"nvidia/Cosmos-Reason2-2B","url":"https://huggingface.co/nvidia/Cosmos-Reason2-2B"}],"capabilities":[{"name":"Physical-AI reasoning VLM","detail":"Spatio-temporal video reasoning, 2D/3D point and box localization, robot planning; 8B beats base Qwen3-VL-8B on robotics (56.90 vs 53.08) and self-driving (67.85 vs 46.38) evals per model card.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/Cosmos-Reason2-8B"},{"name":"Backbone for GR00T N1.7","detail":"Cosmos-Reason2-2B is the VLM backbone of Isaac GR00T N1.7.","first":false,"discovered":"later","source":"https://github.com/NVIDIA/Isaac-GR00T"}],"entry":null,"notes":"Based on Qwen3-VL (8B variant from Qwen3-VL-8B-Instruct, 8.7B params, 32 GB+ GPU). Initial release 2025-12-19, updated 2026-03-10; promoted at CES 2026. Up to 256K input tokens. Its role is folded into Cosmos 3 for new projects.","verified":"2026-09-29","body_md":"Sources: [Cosmos-Reason2-8B card](https://huggingface.co/nvidia/Cosmos-Reason2-8B), [NVIDIA forum: CES 2026 Cosmos announcements](https://forums.developer.nvidia.com/t/nvidia-cosmos-announcements-at-ces-2026/356629).","page_url":"https://postcutoff.com/m/cosmos-reason-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":"https://postcutoff.com/you/cosmos-reason-2/","briefings":null},{"id":"cohere-rerank-4","name":"Cohere Rerank 4 (Pro / Fast)","org":"Cohere","family":"Rerank","released":"2025-12-11","status":"current","type":"embedding","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":32000,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Cohere API","model_id":"rerank-v4.0-pro","endpoint":"https://api.cohere.com/v2/rerank","docs":"https://docs.cohere.com/docs/models"},{"provider":"Cohere API (fast)","model_id":"rerank-v4.0-fast","endpoint":"https://api.cohere.com/v2/rerank"},{"provider":"OpenRouter","model_id":"cohere/rerank-4-pro","url":"https://openrouter.ai/cohere/rerank-4-pro"}],"capabilities":[{"name":"32K-context reranking","detail":"Rerank window grew from 4K (v3.5) to 32K tokens, so whole long documents can be scored.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/models"},{"name":"Pro / Fast tiers","detail":"Two variants: pro for best accuracy, fast for latency-sensitive search.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/models"}],"entry":null,"notes":"Reranker (scores query-document relevance). Previous: rerank-v3.5 (Bedrock cohere.rerank-v3-5:0). Release date from third-party listing; pricing not verified (OpenRouter ~$0.0025/search reported, not checked).","verified":"2026-09-29","body_md":"Second-stage reranking for RAG and enterprise search.\n\n```bash\ncurl https://api.cohere.com/v2/rerank -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"rerank-v4.0-pro\",\"query\":\"capital of France\",\"documents\":[\"Paris is the capital of France.\",\"Berlin is in Germany.\"]}'\n```\n\nSources: https://docs.cohere.com/docs/models","page_url":"https://postcutoff.com/m/cohere-rerank-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":null,"briefings":null},{"id":"fun-cosyvoice3","name":"Fun-CosyVoice3 0.5B (2512) + Fun-ASR-Nano + Fun-Audio-Chat-8B","org":"Alibaba (Tongyi Lab / FunAudioLLM)","family":"FunAudioLLM","released":"2025-12-11","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"FunAudioLLM/Fun-CosyVoice3-0.5B-2512","url":"https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512"},{"provider":"Hugging Face (ASR, 800M)","model_id":"FunAudioLLM/Fun-ASR-Nano-2512","url":"https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512"},{"provider":"Hugging Face (speech chat, 8B)","model_id":"FunAudioLLM/Fun-Audio-Chat-8B","url":"https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B"},{"provider":"GitHub","url":"https://github.com/QwenAudio/CosyVoice"}],"capabilities":[{"name":"Small open multilingual zero-shot TTS","detail":"0.5B model with 9 languages (zh, en, ja, ko, de, es, fr, it, ru) and 18+ Chinese dialects/accents; RL variant reports 0.81% CER / 77.4% speaker similarity (zh) and 1.68% WER / 69.5% similarity (en) on its eval set.","first":false,"discovered":"launch","source":"https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512"},{"name":"Compact far-field ASR (Fun-ASR-Nano, 800M)","detail":"zh/en/ja plus 7 Chinese dialect groups and 26 accents; WER 1.80% AIShell1, 1.76% LibriSpeech-clean; tuned for noisy far-field audio and lyrics over music. MLT-Nano variant covers 31 languages.","first":false,"discovered":"launch","source":"https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512"},{"name":"Open 8B speech chat model with function calling (Fun-Audio-Chat)","detail":"Half-duplex speech-to-speech/speech-to-text LLM (zh/en) with dual-resolution speech representations (5 Hz backbone + 25 Hz head, about 50% less compute); spoken QA, speech function calling, voice empathy.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2512.20156"}],"entry":null,"notes":"HF repo creation dates: CosyVoice3-0.5B-2512 2025-12-11, Fun-ASR-Nano-2512 2025-12-15, Fun-Audio-Chat-8B 2025-12-23. CosyVoice3-0.5B had ~197k downloads in the month to 2026-09-29, one of the most-used open TTS checkpoints. Papers: CosyVoice 3 arXiv 2505.17589, FunAudio-ASR arXiv 2509.12508, Fun-Audio-Chat arXiv 2512.20156. GitHub repo moved from FunAudioLLM/CosyVoice to QwenAudio/CosyVoice. The same Tongyi group's hosted successors are the Qwen-Audio 3.x API models.","verified":"2026-09-29","body_md":"Alibaba Tongyi's open (Apache-2.0) speech stack from December 2025: TTS, ASR and a speech-chat LLM that you can run locally.\n\nSources: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512 · https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512 · https://huggingface.co/FunAudioLLM/Fun-Audio-Chat-8B · https://arxiv.org/abs/2505.17589","page_url":"https://postcutoff.com/m/fun-cosyvoice3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":null,"briefings":null},{"id":"nvidia-magpie-tts-multilingual","name":"NVIDIA MagpieTTS Multilingual 357M","org":"NVIDIA","family":"Nemotron Speech","released":"2025-12-11","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":true,"model_license":"nvidia-open-model-license","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/magpie_tts_multilingual_357m","url":"https://huggingface.co/nvidia/magpie_tts_multilingual_357m"},{"provider":"Hugging Face collection","url":"https://huggingface.co/collections/nvidia/nemotron-speech"}],"capabilities":[{"name":"Small open multilingual TTS for commercial use","detail":"~357-364M-parameter transformer encoder-decoder predicting multi-codebook audio codec tokens; 12 languages (ar, zh, en, fr, de, hi, it, ja, ko, pt, es, vi); 5 built-in English voices; CER 0.34-3.17% across languages per model card; trained on ~54,300 h.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/magpie_tts_multilingual_357m"}],"entry":null,"notes":"Versions: v2512 (HF repo created 2025-12-11), v2602 (Mar 2026), v2607 (2026-07-21); repo last updated 2026-09-09. Zero-shot voice cloning was deliberately removed 'for security reasons'. Part of the Nemotron Speech collection with Parakeet ASR, PersonaPlex and NemotronLabs-VoiceChat.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/nvidia/magpie_tts_multilingual_357m , https://huggingface.co/collections/nvidia/nemotron-speech","page_url":"https://postcutoff.com/m/nvidia-magpie-tts-multilingual/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":null,"briefings":null},{"id":"nova-2-lite","name":"Amazon Nova 2 Lite","org":"Amazon","family":"Nova 2","released":"2025-12-02","status":"current","type":"reasoning-llm","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":64000,"knowledge_cutoff":"2025-10","pricing":{"input":0.3,"output":2.5,"unit":"per 1M tokens (USD) on OpenRouter; Bedrock on-demand price not verified","source":"https://openrouter.ai/amazon/nova-2-lite-v1"},"price_line":"$0.30 in, $2.50 out per 1M tokens","access":[{"provider":"AWS Bedrock","model_id":"amazon.nova-2-lite-v1:0","inference_profiles":["global.amazon.nova-2-lite-v1:0","us.amazon.nova-2-lite-v1:0","eu.amazon.nova-2-lite-v1:0"],"endpoint":"https://bedrock-runtime.{region}.amazonaws.com","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html"},{"provider":"OpenRouter","model_id":"amazon/nova-2-lite-v1","url":"https://openrouter.ai/amazon/nova-2-lite-v1"}],"capabilities":[{"name":"Adjustable extended thinking + 1M context","detail":"Nova 2 generation adds adjustable extended thinking and a 1M-token context for text/image/video input.","first":false,"discovered":"launch","source":"https://www.aboutamazon.com/news/aws/aws-agentic-ai-amazon-bedrock-nova-models"},{"name":"Built-in code interpreter, web grounding, remote MCP","detail":"Nova 2 models support built-in tools (code interpreter, web grounding) and remote MCP tools on Bedrock.","first":false,"discovered":"launch","source":"https://www.aboutamazon.com/news/aws/aws-agentic-ai-amazon-bedrock-nova-models"}],"entry":null,"notes":"Amazon's current GA general model. Nova 2 Pro and Nova 2 Omni were preview-only (Nova Forge) at last check; no Bedrock ids verified.","verified":"2026-09-29","body_md":"Cost-efficient multimodal reasoning model for automation, document processing and support agents.\n\n```bash\naws bedrock-runtime converse --model-id global.amazon.nova-2-lite-v1:0 \\\n  --messages '[{\"role\":\"user\",\"content\":[{\"text\":\"Hello\"}]}]'\n```\n\nSources: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-lite.html , https://docs.aws.amazon.com/nova/ , https://aws.amazon.com/nova/pricing/","page_url":"https://postcutoff.com/m/nova-2-lite/","events_after":826,"major_after":206,"historic_after":41,"missing_at_launch":12,"events_since_release":null,"you_url":"https://postcutoff.com/you/nova-2-lite/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-10-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-10-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-10-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-10.md"}},{"id":"nova-2-sonic","name":"Amazon Nova 2 Sonic","org":"Amazon","family":"Nova 2","released":"2025-12-02","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":64000,"knowledge_cutoff":null,"pricing":{"speech_input":3,"speech_output":12,"text_input":0.33,"text_output":2.75,"unit":"USD per 1M tokens (speech in/out, text in/out); secondary source, not confirmed on the AWS pricing page","source":"https://www.deeplearning.ai/the-batch/nova-2-family-boosts-cost-effective-performance-adds-new-agentic-features"},"price_line":"$0.33 text in, $2.75 text out per 1M tokens","access":[{"provider":"AWS Bedrock","model_id":"amazon.nova-2-sonic-v1:0","endpoint":"https://bedrock-runtime.{region}.amazonaws.com (InvokeModelWithBidirectionalStream)","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html"}],"capabilities":[{"name":"Real-time speech-to-speech","detail":"Single model for natural real-time voice conversations over a bidirectional streaming API (no separate ASR/TTS pipeline).","first":false,"discovered":"launch","source":"https://aws.amazon.com/blogs/aws/introducing-amazon-nova-2-sonic-next-generation-speech-to-speech-model-for-conversational-ai/"},{"name":"1M-token session context","detail":"1M-token context window and 64K max output listed for long-running voice sessions.","first":false,"discovered":"launch","source":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html"},{"name":"Polyglot voices and turn-taking control","detail":"Same voice speaks multiple languages natively (Portuguese and Hindi added vs Nova Sonic); developers set low/medium/high pause sensitivity.","first":false,"discovered":"launch","source":"https://aws.amazon.com/about-aws/whats-new/2025/12/amazon-nova-2-sonic-real-time-conversational-ai"}],"entry":null,"notes":"Technical report (Amazon Nova 2, Dec 2025, https://www.amazon.science/publications/amazon-nova-2-multimodal-reasoning-and-generation-models): Big Bench Audio 87.0 (Artificial Analysis) vs GPT-Realtime (Aug 2025) 83.0 and Gemini 2.5 Flash Live 71.0; BFCL subset 74.5; ComplexFunction 65.2; Common Voice avg WER 6.5 vs 8.4 (GPT-Realtime) across 7 languages; human-preference win rate vs GPT-Realtime above 50% for 6 of 8 voices (e.g. 68.4% Spanish) but 42.4% Hindi and 26.3% Portuguese; vs Gemini 2.5 Flash Live 47.5-77.9%. Comparisons are against 2025 competitors. Successor to Nova Sonic (amazon.nova-sonic-v1:0, Apr 2025). Bedrock only, In-Region in us-east-1, us-west-2, eu-north-1, ap-northeast-1 (no cross-region inference); Standard tier only. Lifecycle Active, EOL no sooner than 2026-12-02. No newer Nova Sonic found as of 2026-09-29; per July 2026 reports Nova 2 Sonic is among the Nova models Amazon keeps developing after its Nova wind-down. Prices from secondary source (AWS Nova pricing page does not list per-token rates).","verified":"2026-09-29","body_md":"Voice agents and conversational IVR on Bedrock.\n\nUse the Bedrock `InvokeModelWithBidirectionalStream` API with model id `amazon.nova-2-sonic-v1:0` (see AWS samples in the model card).\n\nSources: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html · https://cdn.amazon.science/c5/3d/84514a224666b5be6de4b43ef4aa/nova-2-0-technical-report2.pdf","page_url":"https://postcutoff.com/m/nova-2-sonic/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":814,"you_url":null,"briefings":null},{"id":"mistral-large-3","name":"Mistral Large 3","org":"Mistral AI","family":"Mistral Large","released":"2025-12-02","status":"current","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":256000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.5,"output":1.5,"unit":"per 1M tokens (USD)","source":"https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12"},"price_line":"$0.50 in, $1.50 out per 1M tokens","access":[{"provider":"Mistral API","model_id":"mistral-large-2512","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12"},{"provider":"AWS Bedrock","model_id":"mistral.mistral-large-3-675b-instruct","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-mistral-ai-mistral-large-3.html"},{"provider":"OpenRouter","model_id":"mistralai/mistral-large-2512","url":"https://openrouter.ai/mistralai/mistral-large-2512"},{"provider":"Hugging Face","url":"https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512"},{"provider":"Web app","url":"https://chat.mistral.ai"}],"capabilities":[{"name":"675B open-weight MoE under Apache 2.0","detail":"Granular mixture-of-experts with 41B active / 675B total parameters, fully Apache 2.0.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12"},{"name":"Very low price for size","detail":"$0.5 / $1.5 per 1M tokens with 256K context and vision - cheaper than Mistral Medium 3.5.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12"}],"entry":null,"notes":"Alias mistral-large-latest (v25.12). Still GA; for coding/agents Mistral now points to Medium 3.5.","verified":"2026-09-29","body_md":"Large open-weight general-purpose multimodal MoE.\n\n```bash\ncurl https://api.mistral.ai/v1/chat/completions -H \"Authorization: Bearer $MISTRAL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"mistral-large-2512\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12 , https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-mistral-ai-mistral-large-3.html","page_url":"https://postcutoff.com/m/mistral-large-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":814,"you_url":"https://postcutoff.com/you/mistral-large-3/","briefings":null},{"id":"deepseek-v3-2","name":"DeepSeek-V3.2","org":"DeepSeek","family":"DeepSeek V3","released":"2025-12-01","status":"legacy","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"DeepSeek API (retired)","model_id":"deepseek-chat / deepseek-reasoner (no longer serve V3.2)","endpoint":"https://api.deepseek.com","docs":"https://api-docs.deepseek.com/updates"},{"provider":"OpenRouter","model_id":"deepseek/deepseek-v3.2","url":"https://openrouter.ai/deepseek/deepseek-v3.2"},{"provider":"Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2"},{"provider":"Hugging Face (Speciale)","url":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale"}],"capabilities":[{"name":"DeepSeek Sparse Attention (DSA)","detail":"Introduced DSA (first in V3.2-Exp) to cut long-context attention compute while preserving quality.","first":false,"discovered":"launch","source":"https://huggingface.co/deepseek-ai/DeepSeek-V3.2"},{"name":"Hybrid thinking/non-thinking in one model","detail":"deepseek-chat mapped to non-thinking mode and deepseek-reasoner to thinking mode of the same V3.2 weights.","first":false,"discovered":"launch","source":"https://api-docs.deepseek.com/updates"},{"name":"V3.2-Speciale reasoning variant","detail":"Separate high-compute Speciale variant served briefly on a temporary endpoint (no tool calls) until 2025-12-15; weights released.","first":false,"discovered":"launch","source":"https://api-docs.deepseek.com/updates"}],"entry":null,"notes":"API aliases deepseek-chat/deepseek-reasoner moved to V4-Flash on 2026-04-24 and were scheduled for discontinuation on 2026-07-24; V3.2 now only via open weights/third parties. Pricing not verified (no first-party price).","verified":"2026-09-29","body_md":"Open-weight (MIT) predecessor of V4; still widely self-hosted and on third-party APIs. Use deepseek-v4-pro / deepseek-flash on the first-party API instead.\n\nSources: https://api-docs.deepseek.com/updates · https://huggingface.co/deepseek-ai/DeepSeek-V3.2 · https://openrouter.ai/deepseek/deepseek-v3.2","page_url":"https://postcutoff.com/m/deepseek-v3-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":814,"you_url":null,"briefings":null},{"id":"runway-gen-4-5","name":"Runway Gen-4.5","org":"Runway","family":"Gen-4","released":"2025-12-01","status":"current","type":"video-gen","modality_in":["text","image"],"modality_out":["video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.12,"unit":"per second of video (12 credits/s at $0.01/credit; pro/HDR formats extra)","source":"https://docs.dev.runwayml.com/guides/pricing/"},"price_line":"$0.12 per second","access":[{"provider":"Runway API","model_id":"gen4.5","endpoint":"https://api.dev.runwayml.com/v1/image_to_video","docs":"https://docs.dev.runwayml.com/guides/models/"},{"provider":"Web app","url":"https://app.runwayml.com"}],"capabilities":[{"name":"#1 on Artificial Analysis text-to-video at launch","detail":"Launched as the top model on the Artificial Analysis Text-to-Video leaderboard (1,247 Elo), with better physics (liquids, momentum, collisions).","first":false,"discovered":"launch","source":"https://runway.com/research/introducing-runway-gen-4.5"},{"name":"HDR and professional output formats","detail":"API can output ProRes, PNG/EXR sequences, 10-bit SDR and HDR10/HLG/ACEScg masters (Gen-4.5 only for HDR).","first":false,"discovered":"later","source":"https://docs.dev.runwayml.com/guides/models/"}],"entry":null,"notes":"Announced 2025-12-01; added to Runway API 2026-02-10 (text-to-video and image-to-video, 2-10 s). Cheaper sibling gen4_turbo (5 credits/s). gen4_aleph and gen3a_turbo removed from API 2026-07-30. Requires header X-Runway-Version: 2024-11-06.","verified":"2026-09-29","body_md":"Runway's flagship text/image-to-video model. Runway's API also resells third-party models (Veo, Seedance, etc.).\n\n```bash\ncurl -X POST https://api.dev.runwayml.com/v1/image_to_video -H \"Authorization: Bearer $RUNWAYML_API_SECRET\" \\\n  -H \"X-Runway-Version: 2024-11-06\" -H \"Content-Type: application/json\" \\\n  -d '{\"model\":\"gen4.5\",\"promptText\":\"a drone shot over a misty fjord\",\"ratio\":\"1280:720\",\"duration\":5}'\n# poll GET https://api.dev.runwayml.com/v1/tasks/{id}\n```\n\nSources: https://docs.dev.runwayml.com/guides/models/ , https://docs.dev.runwayml.com/guides/pricing/ , https://docs.dev.runwayml.com/api-details/api_changelog/ , https://runway.com/research/introducing-runway-gen-4.5","page_url":"https://postcutoff.com/m/runway-gen-4-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":814,"you_url":null,"briefings":null},{"id":"flux-2-max","name":"FLUX.2 [max]","org":"Black Forest Labs","family":"FLUX.2","released":"2025-12","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.07,"unit":"per image from (text-to-image and editing; scales with megapixels)","source":"https://docs.bfl.ai/quick_start/pricing"},"price_line":"$0.07 per image","access":[{"provider":"BFL API","model_id":"flux-2-max","endpoint":"https://api.bfl.ai/v1/flux-2-max","docs":"https://docs.bfl.ai/flux_2/flux2_overview"},{"provider":"Web app","url":"https://playground.bfl.ai"}],"capabilities":[{"name":"Grounded generation with web search","detail":"Can pull real-time web context (grounding search) into generations, e.g. current events or real products.","first":false,"discovered":"launch","source":"https://bfl.ai/models/flux-2-max"},{"name":"Highest editing consistency in FLUX.2","detail":"Top FLUX.2 tier for prompt following, style fidelity, character consistency and retexturing/product photography.","first":false,"discovered":"launch","source":"https://bfl.ai/models/flux-2-max"}],"entry":null,"notes":"Release month (Dec 2025) not confirmed on an official page. Endpoint confirmed in https://api.bfl.ai/openapi.json.","verified":"2026-09-29","body_md":"Highest-quality FLUX.2 variant, with web-grounded generation. Same async pattern as flux-2-pro (POST, then poll `polling_url`).\n\n```bash\ncurl -X POST https://api.bfl.ai/v1/flux-2-max -H \"x-key: $BFL_API_KEY\" -H \"Content-Type: application/json\" -d '{\"prompt\":\"...\"}'\n```\n\nSources: https://bfl.ai/models/flux-2-max , https://docs.bfl.ai/quick_start/pricing","page_url":"https://postcutoff.com/m/flux-2-max/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":null,"briefings":null},{"id":"glm-4-6v","name":"GLM-4.6V","org":"Zhipu AI (Z.ai)","family":"GLM-4","released":"2025-12","status":"legacy","type":"multimodal","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":128000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":0.9,"cache_read":0.05,"unit":"per 1M tokens (USD); GLM-4.6V-FlashX 0.04/0.4; GLM-4.6V-Flash free","source":"https://docs.z.ai/guides/overview/pricing"},"price_line":"$0.30 in, $0.90 out per 1M tokens","access":[{"provider":"Z.ai API","model_id":"glm-4.6v","endpoint":"https://api.z.ai/api/paas/v4","docs":"https://docs.z.ai/guides/overview/pricing"},{"provider":"OpenRouter","model_id":"z-ai/glm-4.6v","url":"https://openrouter.ai/z-ai/glm-4.6v"},{"provider":"Hugging Face","url":"https://huggingface.co/zai-org/GLM-4.6V"},{"provider":"Hugging Face (Flash)","url":"https://huggingface.co/zai-org/GLM-4.6V-Flash"},{"provider":"Web app","url":"https://chat.z.ai"}],"capabilities":[{"name":"Native multimodal function calling","detail":"First GLM vision model with native function calling (images can be passed to and returned from tools).","first":false,"discovered":"launch","source":"https://huggingface.co/zai-org/GLM-4.6V"},{"name":"Interleaved image-text generation","detail":"Builds mixed image-text content from documents and tool-retrieved images; also frontend replication from screenshots.","first":false,"discovered":"launch","source":"https://huggingface.co/zai-org/GLM-4.6V"}],"entry":null,"notes":"Still sold on Z.ai (with FlashX and free Flash variants) but superseded by the natively multimodal GLM-5.3-Flash. Model id casing on Z.ai assumed lowercase glm-4.6v (listed as GLM-4.6V). Release day not verified (HF 2025-12-07).","verified":null,"body_md":"Open-weight (MIT) vision-language GLM; cheapest/free Z.ai vision option via GLM-4.6V-Flash.\n\nSources: https://huggingface.co/zai-org/GLM-4.6V · https://docs.z.ai/guides/overview/pricing","page_url":"https://postcutoff.com/m/glm-4-6v/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":809,"you_url":null,"briefings":null},{"id":"deepseekmath-v2","name":"DeepSeekMath-V2","org":"DeepSeek","family":"DeepSeekMath","released":"2025-11-27","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/deepseek-ai/DeepSeek-Math-V2"}],"capabilities":[{"name":"Self-verifiable proof generation","detail":"Generator trained against an LLM proof verifier and meta-verifier; reached IMO 2025 / CMO 2024 gold level and 118/120 on Putnam 2024 with scaled test-time compute.","first":true,"discovered":"launch","source":"https://arxiv.org/abs/2511.22570"}],"entry":"2025-11-27-deepseekmath-v2","notes":"685B open-weights (Apache 2.0) math prover built on DeepSeek-V3.2-Exp-Base; inference uses the DeepSeek-V3.2-Exp code. No first-party API endpoint verified. 'first' = first open-weights model at IMO-gold level (per the paper's claims).","verified":"2026-09-29","body_md":"Download the weights from Hugging Face and serve them with the DeepSeek-V3.2-Exp inference stack. Source: https://huggingface.co/deepseek-ai/DeepSeek-Math-V2 · paper https://arxiv.org/abs/2511.22570","page_url":"https://postcutoff.com/m/deepseekmath-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":816,"you_url":"https://postcutoff.com/you/deepseekmath-v2/","briefings":null},{"id":"flux-2-dev","name":"FLUX.2 [dev]","org":"Black Forest Labs","family":"FLUX.2","released":"2025-11-25","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":true,"model_license":"FLUX Non-Commercial License","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/black-forest-labs/FLUX.2-dev"},{"provider":"Hugging Face (NVFP4)","url":"https://huggingface.co/black-forest-labs/FLUX.2-dev-NVFP4"}],"capabilities":[{"name":"32B open-weight generation + multi-reference editing","detail":"32B open-weight model doing text-to-image, single- and multi-reference editing in one checkpoint; BFL claims it beats all open-weight alternatives.","first":false,"discovered":"launch","source":"https://bfl.ai/blog/flux-2"},{"name":"VLM-conditioned rectified flow transformer","detail":"Pairs a Mistral-3 24B vision-language model with a rectified flow transformer for world knowledge and prompt understanding.","first":false,"discovered":"launch","source":"https://bfl.ai/blog/flux-2"}],"entry":null,"notes":"Open weights only (no /v1/flux-2-dev endpoint in BFL API openapi.json); commercial use needs a BFL license (https://bfl.ai/licensing). Hosted by many third parties. Pricing n/a.","verified":"2026-09-29","body_md":"Best open-weight FLUX.2 checkpoint for local/self-hosted image generation and editing (large VRAM needs; use FP8/NVFP4 variants on consumer GPUs).\n\n```python\nfrom diffusers import Flux2Pipeline\npipe = Flux2Pipeline.from_pretrained(\"black-forest-labs/FLUX.2-dev\").to(\"cuda\")\n```\n\nSources: https://bfl.ai/blog/flux-2 , https://huggingface.co/black-forest-labs/FLUX.2-dev","page_url":"https://postcutoff.com/m/flux-2-dev/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":818,"you_url":null,"briefings":null},{"id":"flux-2-pro","name":"FLUX.2 [pro]","org":"Black Forest Labs","family":"FLUX.2","released":"2025-11-25","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.03,"unit":"per image from (1MP text-to-image; editing from 0.045; scales with megapixels)","source":"https://docs.bfl.ai/quick_start/pricing"},"price_line":"$0.03 per image","access":[{"provider":"BFL API","model_id":"flux-2-pro","endpoint":"https://api.bfl.ai/v1/flux-2-pro","docs":"https://docs.bfl.ai/flux_2/flux2_overview"},{"provider":"Web app","url":"https://playground.bfl.ai"}],"capabilities":[{"name":"Multi-reference editing (up to 10 images)","detail":"Generates and edits with up to 10 reference images for character/product/style consistency, in one model with text-to-image.","first":false,"discovered":"launch","source":"https://bfl.ai/blog/flux-2"},{"name":"4MP editing and production-grade typography","detail":"Image editing up to 4 megapixels; reliable fine text for infographics, memes and UI mockups.","first":false,"discovered":"launch","source":"https://bfl.ai/blog/flux-2"}],"entry":null,"notes":"flux-2-pro is a fixed snapshot; flux-2-pro-preview tracks the latest [pro]. Siblings: flux-2-flex (from $0.05, step/guidance control), flux-2-max. Uses Mistral-3 24B VLM + rectified flow transformer.","verified":"2026-09-29","body_md":"BFL's production workhorse for text-to-image and multi-reference editing.\n\n```bash\ncurl -X POST https://api.bfl.ai/v1/flux-2-pro -H \"x-key: $BFL_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"prompt\":\"product shot of a ceramic mug on linen, soft window light\"}'\n# response: {id, polling_url} -> GET polling_url until status Ready\n```\n\nSources: https://docs.bfl.ai/flux_2/flux2_overview , https://docs.bfl.ai/quick_start/pricing , https://bfl.ai/blog/flux-2","page_url":"https://postcutoff.com/m/flux-2-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":818,"you_url":null,"briefings":null},{"id":"claude-opus-4-5","name":"Claude Opus 4.5","org":"Anthropic","family":"Claude 4","released":"2025-11-24","status":"legacy","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":64000,"knowledge_cutoff":"2025-05","pricing":{"input":5,"output":25,"cache_read":0.5,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$5 in, $25 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-4-5-20251101","alias":"claude-opus-4-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/opus-4-5/overview"},{"provider":"AWS Bedrock (InvokeModel)","model_id":"anthropic.claude-opus-4-5-20251101-v1:0","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy"},{"provider":"Google Cloud Vertex AI","model_id":"claude-opus-4-5@20251101","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-opus-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-opus-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-4.5","url":"https://openrouter.ai/anthropic/claude-opus-4.5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Beat all human candidates on Anthropic's engineering exam","detail":"Scored higher than any human candidate on Anthropic's take-home engineering exam within the 2-hour limit.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-5"},{"name":"Effort parameter","detail":"First model with the effort parameter. At medium effort it matched Sonnet 4.5's best score with 76% fewer output tokens.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-5"},{"name":"Prompt-injection robustness","detail":"Anthropic claimed it was harder to trick with prompt injection than any other frontier model at the time.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-5"},{"name":"Opus price cut","detail":"Opus-class pricing dropped to $5/$25 per MTok, from $15/$75.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-5"}],"entry":null,"notes":"Snapshot claude-opus-4-5-20251101 (alias claude-opus-4-5). Extended thinking (budget_tokens); effort low/medium/high. Training data cutoff Aug 2025. Retirement not sooner than 2026-11-24.","verified":"2026-09-29","body_md":"Legacy Opus; migrate to `claude-opus-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-opus-4-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/opus-4-5/overview\n- Announcement: https://www.anthropic.com/news/claude-opus-4-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-opus-4-5/","events_after":863,"major_after":219,"historic_after":44,"missing_at_launch":44,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-opus-4-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-05-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-05-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-05-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-05.md"}},{"id":"gemini-3-pro-image","name":"Nano Banana Pro (Gemini 3 Pro Image)","org":"Google DeepMind","family":"Gemini Image","released":"2025-11-20","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image","text"],"open_weights":false,"model_license":"proprietary","context_window":65536,"max_output":32768,"knowledge_cutoff":null,"pricing":{"input":2,"output":12,"per_image_1k_2k":0.134,"per_image_4k":0.24,"unit":"per 1M tokens text/image input and text output; image output $120 per 1M tokens = $0.134 per 1K/2K image, $0.24 per 4K image","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$2 in, $12 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-3-pro-image","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-image"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-pro-image"},{"provider":"OpenRouter","model_id":"google/gemini-3-pro-image"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Accurate multilingual text in images","detail":"Correct, legible text rendering in many languages, fonts and calligraphy; suited to infographics and mockups.","first":false,"discovered":"launch","source":"https://blog.google/technology/ai/nano-banana-pro/"},{"name":"Search-grounded visuals","detail":"Uses Google Search to visualize real-time info (weather, sports, recipes) and factual data visualizations.","first":false,"discovered":"launch","source":"https://blog.google/technology/ai/nano-banana-pro/"},{"name":"Multi-image composition","detail":"Blends up to 14 images while keeping resemblance of up to 5 people; up to 4K with lighting/depth-of-field edits.","first":false,"discovered":"launch","source":"https://blog.google/technology/ai/nano-banana-pro/"}],"entry":null,"notes":"Launched 2025-11-20 as gemini-3-pro-image-preview (still on OpenRouter); stable id GA 2026-05-28. Highest-quality but priciest Gemini image model.","verified":"2026-09-29","body_md":"Studio-quality image generation built on Gemini 3 Pro reasoning; best for complex graphic design and text-heavy images.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Infographic explaining photosynthesis, labeled\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-image), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [launch blog](https://blog.google/technology/ai/nano-banana-pro/).","page_url":"https://postcutoff.com/m/gemini-3-pro-image/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":819,"you_url":null,"briefings":null},{"id":"dia2","name":"Nari Labs Dia2 (1B / 2B)","org":"Nari Labs","family":"Dia","released":"2025-11-19","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nari-labs/Dia2-2B","url":"https://huggingface.co/nari-labs/Dia2-2B"},{"provider":"GitHub","url":"https://github.com/nari-labs/dia2"}],"capabilities":[{"name":"Streaming multi-speaker dialogue TTS","detail":"Generates [S1]/[S2] dialogue and starts producing audio from the first few input tokens (no need for full text); conditions on audio prefixes for real-time conversation; up to ~2 min per generation (Mimi codec, 12.5 Hz); word-level timestamps.","first":false,"discovered":"launch","source":"https://huggingface.co/nari-labs/Dia2-2B"}],"entry":null,"notes":"English only. Successor to Dia-1.6B (April 2025, github.com/nari-labs/dia). Release date 2025-11-19 from secondary sources (GitHub releases page).","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/nari-labs/Dia2-2B , https://github.com/nari-labs/dia2","page_url":"https://postcutoff.com/m/dia2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":820,"you_url":null,"briefings":null},{"id":"pi-0-6","name":"π0.6 / π*0.6","org":"Physical Intelligence","family":"π (pi)","released":"2025-11-17","status":"legacy","type":"robotics","modality_in":["text","image"],"modality_out":["action","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"None (internal to Physical Intelligence; model card and paper only)","url":"https://www.pi.website/blog/pistar06","docs":"https://website.pi-asset.com/pi06star/PI06_model_card.pdf"}],"capabilities":[{"name":"Recap - RL from real-world experience and corrections","detail":"π*0.6 improves π0.6 with Recap (RL with Experience & Corrections via Advantage-conditioned Policies): demonstrations, then human interventions, then autonomous-trial RL; over 2x throughput and roughly halved failure rates on hard tasks.","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pistar06"},{"name":"Hours-long autonomous operation","detail":"Made espresso drinks for 18 hours straight, folded 50 novel laundry items in a new home, and assembled/labeled 59 factory boxes.","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pistar06"}],"entry":"2025-11-17-physical-intelligence-pi-star-0-6-recap","notes":"π0.6: ~5B-parameter VLA with a Gemma 3 4B backbone and ~860M-parameter action expert, keeps π0.5's hierarchical design (per model card, 2025-11-17, via search snippet). No weights or API. Superseded by π0.7 (2026-04). pi.website blocked automated fetches on 2026-09-29; details taken from search snippets of the blog/model card.","verified":null,"body_md":"Changelog: 2026-09-29 linked new entry 2025-11-17-physical-intelligence-pi-star-0-6-recap and arXiv 2511.14759.\n\nSources: [π*0.6 blog](https://www.pi.website/blog/pistar06), [paper PDF](https://www.pi.website/download/pistar06.pdf), [π0.6 model card](https://website.pi-asset.com/pi06star/PI06_model_card.pdf), [Humanoids Daily](https://www.humanoidsdaily.com/news/physical-intelligence-claims-rl-is-back-with-new-model-that-learns-from-its-own-mistakes).","page_url":"https://postcutoff.com/m/pi-0-6/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":821,"you_url":null,"briefings":null},{"id":"elevenlabs-scribe-v2-realtime","name":"Scribe v2 Realtime","org":"ElevenLabs","family":"Scribe","released":"2025-11-11","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_hour":0.39,"unit":"per hour of audio (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.39 per hour","access":[{"provider":"ElevenLabs API (WebSocket)","model_id":"scribe_v2_realtime","endpoint":"wss://api.elevenlabs.io/v1/speech-to-text/realtime","docs":"https://elevenlabs.io/docs/models"}],"capabilities":[{"name":"~150 ms streaming STT with next-word prediction","detail":"Under 150 ms transcription latency with 'negative latency' next-word and punctuation prediction; VAD, manual commit, mid-conversation language switching; 90+ languages; PCM 48 kHz and u-law.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/introducing-scribe-v2-realtime"},{"name":"Realtime entity detection","detail":"Entity detection added to realtime transcription on 2026-08-03.","first":false,"discovered":"later","source":"https://elevenlabs.io/docs/changelog"}],"entry":null,"notes":"Launched 2025-11-11; claims 93.5% accuracy across 30 European and Asian languages (company figure). EU and India data residency, zero-retention mode.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/blog/introducing-scribe-v2-realtime","page_url":"https://postcutoff.com/m/elevenlabs-scribe-v2-realtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":822,"you_url":null,"briefings":null},{"id":"omnilingual-asr","name":"Omnilingual ASR","org":"Meta","family":"Omnilingual ASR","released":"2025-11-10","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"GitHub (fairseq2 checkpoints)","model_id":"omniASR_LLM_7B_v2","url":"https://github.com/facebookresearch/omnilingual-asr"},{"provider":"Hugging Face (demo space and dataset)","url":"https://huggingface.co/facebook"}],"capabilities":[{"name":"ASR for 1,600+ languages","detail":"Transcribes 1,600+ languages, ~500 of them never before supported by any ASR system (Whisper covers 99).","first":true,"discovered":"launch","source":"https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/"},{"name":"Zero-shot in-context language extension","detail":"omniASR_LLM_7B_ZS transcribes new languages from a few paired audio-text examples at inference, extending potential coverage to 5,400+ languages.","first":false,"discovered":"launch","source":"https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/"}],"entry":null,"notes":"Open (Apache 2.0) suite: CTC and LLM-ASR models at 300M/1B/3B/7B, v2 checkpoints and 'Unlimited' long-audio LLM-ASR variants added December 2025, plus a 7B wav2vec 2.0 speech encoder and a corpus covering 350+ underserved languages. Checkpoints download via fairseq2 (e.g. https://dl.fbaipublicfiles.com/mms/omniASR-LLM-7B-v2.pt). Successor to MMS. The 'first' claim is Meta's ('never previously supported by any ASR model').","verified":"2026-09-29","body_md":"Meta's open massively multilingual speech recognition (Nov 2025, still Meta's current open ASR).\n\nSources: https://github.com/facebookresearch/omnilingual-asr · https://ai.meta.com/blog/omnilingual-asr-advancing-automatic-speech-recognition/","page_url":"https://postcutoff.com/m/omnilingual-asr/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":823,"you_url":null,"briefings":null},{"id":"step-audio-editx","name":"StepFun Step-Audio-EditX","org":"StepFun","family":"Step-Audio","released":"2025-11-06","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0 (code; check model card for weights)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"stepfun-ai/Step-Audio-EditX","url":"https://huggingface.co/stepfun-ai/Step-Audio-EditX"},{"provider":"Hugging Face (4-bit)","model_id":"stepfun-ai/Step-Audio-EditX-AWQ-4bit","url":"https://huggingface.co/stepfun-ai/Step-Audio-EditX-AWQ-4bit"},{"provider":"GitHub","url":"https://github.com/stepfun-ai/Step-Audio-EditX"}],"capabilities":[{"name":"Iterative LLM-based audio editing","detail":"3B RL-trained audio LLM that edits emotion, speaking style and paralinguistics of existing speech step by step, plus zero-shot TTS cloning (Mandarin, English, Sichuanese, Cantonese; Japanese/Korean added 2025-11-28).","first":false,"discovered":"launch","source":"https://github.com/stepfun-ai/Step-Audio-EditX"}],"entry":null,"notes":"Official changelog lists a new model release on 2026-01-29 (overall ~4% improvement; new paralinguistic tags such as exhale, inhale, chuckle, clears throat, giggle; SFT/DPO/GRPO training code released); HF weights updated 2026-01-23/24, README edits to 2026-02-14. No March 2026 release appears in the official GitHub/HF changelog, so Artificial Analysis's 'Step Audio EditX (Mar 2026)' label (#3 open weights, ~1095 Elo, Sept 2026) probably refers to the Jan 2026 weights or a hosted snapshot (unverified).","verified":"2026-09-29","body_md":"Sources: https://github.com/stepfun-ai/Step-Audio-EditX , https://huggingface.co/stepfun-ai/Step-Audio-EditX","page_url":"https://postcutoff.com/m/step-audio-editx/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":823,"you_url":null,"briefings":null},{"id":"kimi-k2-thinking","name":"Kimi K2 Thinking","org":"Moonshot AI","family":"Kimi K2","released":"2025-11","status":"retired","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"modified-mit","context_window":262144,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Kimi API (discontinued)","model_id":"kimi-k2-thinking / kimi-k2-thinking-turbo","endpoint":"https://api.moonshot.ai/v1","docs":"https://platform.kimi.ai/docs/models"},{"provider":"OpenRouter","model_id":"moonshotai/kimi-k2-thinking","url":"https://openrouter.ai/moonshotai/kimi-k2-thinking"},{"provider":"Hugging Face","url":"https://huggingface.co/moonshotai/Kimi-K2-Thinking"}],"capabilities":[{"name":"Long tool-call chains","detail":"Interleaves reasoning with function calls and stays coherent across 200-300 sequential tool calls (vs 30-50 for earlier models, per Moonshot).","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2-Thinking"},{"name":"Native INT4 via quantization-aware training","detail":"QAT in post-training gives a lossless ~2x speed-up at INT4 on a 1T/32B-active MoE.","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2-Thinking"},{"name":"Heavy mode","detail":"Parallel 8-trajectory rollout with reflective aggregation used for top benchmark results (HLE, BrowseComp).","first":false,"discovered":"launch","source":"https://huggingface.co/moonshotai/Kimi-K2-Thinking"}],"entry":null,"notes":"kimi-k2 series (incl. K2 Thinking, K2-0905, K2-0711) discontinued on the Kimi API on 2026-05-25; Moonshot recommends kimi-k3. Still available as open weights and via third parties. Pricing not verified.","verified":"2026-09-29","body_md":"Landmark open-weight thinking agent (Nov 2025). Use for self-hosting or research; for API use kimi-k3 or kimi-k2.6.\n\nSources: https://huggingface.co/moonshotai/Kimi-K2-Thinking · https://platform.kimi.ai/docs/models","page_url":"https://postcutoff.com/m/kimi-k2-thinking/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":816,"you_url":null,"briefings":null},{"id":"nova-premier","name":"Amazon Nova Premier","org":"Amazon","family":"Nova","released":"2025-10-31","status":"retired","type":"multimodal","modality_in":["text","image","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1000000,"max_output":25000,"knowledge_cutoff":"2024-10","pricing":{"input":2.5,"output":12.5,"unit":"per 1M tokens (USD) on OpenRouter","source":"https://openrouter.ai/amazon/nova-premier-v1"},"price_line":"$2.50 in, $12.50 out per 1M tokens","access":[{"provider":"AWS Bedrock","model_id":"amazon.nova-premier-v1:0","inference_profiles":["us.amazon.nova-premier-v1:0"],"docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html"},{"provider":"OpenRouter","model_id":"amazon/nova-premier-v1","url":"https://openrouter.ai/amazon/nova-premier-v1"}],"capabilities":[{"name":"Teacher model for distillation","detail":"Positioned for complex reasoning, agentic workflows and as a teacher for Bedrock model distillation into smaller Nova models.","first":false,"discovered":"launch","source":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html"},{"name":"1M context multimodal reasoning","detail":"1M-token context over text, image and video with reasoning support - largest first-gen Nova.","first":false,"discovered":"launch","source":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html"}],"entry":null,"notes":"Bedrock card shows lifecycle Legacy with EOL date 2026-09-14 (passed); may still be listed. Use Nova 2 Lite instead. Launch date as shown on Bedrock card. Nova Pro/Lite/Micro (v1) and Nova Canvas/Reel (EOL 2026-09-30) are also legacy.","verified":"2026-09-29","body_md":"Previous-generation top Nova model; migrate to Nova 2.\n\nSources: https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-premier.html , https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards.html","page_url":"https://postcutoff.com/m/nova-premier/","events_after":893,"major_after":234,"historic_after":49,"missing_at_launch":67,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-10-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-10-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-10-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-10.md"}},{"id":"lyria-2","name":"Lyria 2","org":"Google DeepMind","family":"Lyria","released":"2025-10-27","status":"legacy","type":"music","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_clip":0.06,"unit":"USD per generation (one ~30 s clip) on Vertex AI","source":"https://cloud.google.com/vertex-ai/generative-ai/pricing"},"price_line":"$0.06 per clip","access":[{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"lyria-002","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-002"}],"capabilities":[{"name":"Instrumental clips with negative prompting","detail":"Text-to-music instrumental clips up to 32.8 s, 48 kHz WAV, up to 4 clips per prompt, negative prompts supported; US English prompts only.","first":false,"discovered":"launch","source":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/lyria/lyria-002"}],"entry":null,"notes":"Vertex page lists lyria-002 as GA with release date 2025-10-27 (Lyria 2 was first shown publicly in 2025; earlier preview dates not re-verified). No vocals, lyrics or image input; superseded by Lyria 3 / 3.5 but still GA on Vertex, global region only. Status 'legacy' is our judgement (no deprecation announced).","verified":"2026-09-29","body_md":"Older instrumental-only Lyria on Vertex; use Lyria 3.5 for songs with vocals.","page_url":"https://postcutoff.com/m/lyria-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":828,"you_url":null,"briefings":null},{"id":"claude-haiku-4-5","name":"Claude Haiku 4.5","org":"Anthropic","family":"Claude 4","released":"2025-10-15","status":"current","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":64000,"knowledge_cutoff":"2025-02","pricing":{"input":1,"output":5,"cache_read":0.1,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$1 in, $5 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-haiku-4-5-20251001","alias":"claude-haiku-4-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/haiku-4-5/overview"},{"provider":"AWS Bedrock (Messages API / Mantle)","model_id":"anthropic.claude-haiku-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock"},{"provider":"AWS Bedrock (InvokeModel)","model_id":"anthropic.claude-haiku-4-5-20251001-v1:0","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy"},{"provider":"Google Cloud Vertex AI","model_id":"claude-haiku-4-5@20251001","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-haiku-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-haiku-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-haiku-4.5","url":"https://openrouter.ai/anthropic/claude-haiku-4.5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"Sonnet-4-class coding at Haiku price","detail":"73.3% on SWE-bench Verified, roughly matching Sonnet 4 at one-third the cost and over 2x the speed.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-haiku-4-5"},{"name":"Sub-agent workhorse","detail":"Reaches about 90% of Sonnet 4.5 on Augment's agentic eval; Anthropic positions it for multi-agent orchestration.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-haiku-4-5"},{"name":"Haiku with extended thinking and computer use","detail":"First Haiku model with extended thinking; it also surpasses Sonnet 4 on some computer-use tasks.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-haiku-4-5"}],"entry":null,"notes":"Fastest/cheapest current Claude. Snapshot claude-haiku-4-5-20251001 (alias claude-haiku-4-5). Uses extended thinking (thinking type enabled + budget_tokens), no effort parameter. Training data cutoff Jul 2025. Retirement not sooner than 2026-10-15. Batch $0.50/$2.50.","verified":"2026-09-29","body_md":"Cheap, low-latency model for classification, extraction, sub-agents and high-volume chat.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-haiku-4-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/haiku-4-5/overview\n- Announcement: https://www.anthropic.com/news/claude-haiku-4-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-haiku-4-5/","events_after":874,"major_after":225,"historic_after":45,"missing_at_launch":40,"events_since_release":null,"you_url":"https://postcutoff.com/you/claude-haiku-4-5/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-02-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-02-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-02-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-02.md"}},{"id":"veo-3-1","name":"Veo 3.1","org":"Google DeepMind","family":"Veo","released":"2025-10-15","status":"preview","type":"video-gen","modality_in":["text","image","video"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.4,"per_second_4k":0.6,"unit":"per second of video (veo-3.1-generate-preview, 720p/1080p $0.40, 4K $0.60). Fast: $0.10 (720p) / $0.12 (1080p) / $0.30 (4K). Lite: $0.05 (720p) / $0.08 (1080p)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.40 per second","access":[{"provider":"Gemini API","model_id":"veo-3.1-generate-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning","docs":"https://ai.google.dev/gemini-api/docs/veo"},{"provider":"Gemini API","model_id":"veo-3.1-fast-generate-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-fast-generate-preview:predictLongRunning"},{"provider":"Gemini API","model_id":"veo-3.1-lite-generate-preview","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-lite-generate-preview:predictLongRunning"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/veo/3-1-generate"},{"provider":"Web app","url":"https://gemini.google.com"}],"capabilities":[{"name":"Native audio in every clip","detail":"Dialogue, SFX and ambience generated with the video; Veo 3.1 extended audio to Ingredients-to-Video, Frames-to-Video and Extend.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/products/veo-updates-flow/"},{"name":"Reference images and first/last frame control","detail":"Multiple reference images for character/object/style consistency; generate a bridge between a start and end frame.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/products/veo-updates-flow/"},{"name":"Video extension to a minute+","detail":"Extend clips (720p) to build longer continuous scenes.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/veo"}],"entry":null,"notes":"Specs: 4/6/8 s, 720p/1080p/4K (1080p/4K need 8 s; no 4K on Lite), 16:9 or 9:16, 24 fps. Standard + Fast released 2025-10-15, Lite 2026-03-31; all still preview ids in the Gemini API. Veo 2 and Veo 3.0 sunset 2026-06-30. Google now recommends Gemini Omni Flash as default video model.","verified":"2026-09-29","body_md":"Google's dedicated text/image-to-video model with native audio, served as long-running operations.\n\n```bash\ncurl -X POST \"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"instances\":[{\"prompt\":\"A dramatic sunset over mountains\"}]}'\n# then poll the returned operation name\n```\n\nSources: [Veo guide](https://ai.google.dev/gemini-api/docs/veo), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [Flow blog](https://blog.google/innovation-and-ai/products/veo-updates-flow/).","page_url":"https://postcutoff.com/m/veo-3-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":833,"you_url":null,"briefings":null},{"id":"chirp-3","name":"Chirp 3 Transcription (Google Cloud Speech-to-Text)","org":"Google","family":"Chirp 3","released":"2025-10-13","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.016,"unit":"USD per minute, Speech-to-Text V2 standard recognition, first 500K min/month (tiers down to $0.004/min above 2M min)","source":"https://cloud.google.com/speech-to-text/pricing"},"price_line":"$0.016 per minute","access":[{"provider":"Google Cloud Speech-to-Text API V2","model_id":"chirp_3","endpoint":"https://speech.googleapis.com/v2 (Recognize, StreamingRecognize, BatchRecognize)","docs":"https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3"}],"capabilities":[{"name":"Multilingual ASR with language-agnostic mode","detail":"~100+ languages/locales (about 20 GA), language_codes=['auto'] for language-agnostic transcription, diarization in ~15 languages, speech adaptation.","first":false,"discovered":"launch","source":"https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3"}],"entry":null,"notes":"Private preview 2025-04-11, public preview 2025-08-29, GA 2025-10-13 (US/EU multi-region). No word-level timestamps or word confidence. For developers, Gemini 3.5 Transcribe (2026-08-26) claims 70% faster time-to-final than Chirp 3.","verified":"2026-09-29","body_md":"Sources: [Chirp 3 docs](https://docs.cloud.google.com/speech-to-text/docs/models/chirp-3), [STT release notes](https://docs.cloud.google.com/speech-to-text/docs/release-notes), [pricing](https://cloud.google.com/speech-to-text/pricing).","page_url":"https://postcutoff.com/m/chirp-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"cosmos-predict-2-5","name":"Cosmos Predict 2.5 / Transfer 2.5","org":"NVIDIA","family":"Cosmos","released":"2025-10-06","status":"legacy","type":"world-model","modality_in":["text","image","video"],"modality_out":["video"],"open_weights":true,"model_license":"NVIDIA Open Model License (commercial use allowed)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/Cosmos-Predict2.5-2B","url":"https://huggingface.co/nvidia/Cosmos-Predict2.5-2B"},{"provider":"Hugging Face","model_id":"nvidia/Cosmos-Predict2.5-14B","url":"https://huggingface.co/nvidia/Cosmos-Predict2.5-14B"},{"provider":"Hugging Face","model_id":"nvidia/Cosmos-Transfer2.5-2B","url":"https://huggingface.co/nvidia/Cosmos-Transfer2.5-2B"},{"provider":"GitHub","url":"https://github.com/nvidia-cosmos/cosmos-transfer2.5"}],"capabilities":[{"name":"Unified Text2World / Image2World / Video2World","detail":"Single diffusion transformer for physics-aware video world generation (720p, 16 fps, ~5 s clips) for robotics and AV synthetic data.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/Cosmos-Predict2.5-2B"},{"name":"Multi-control world-to-world transfer","detail":"Transfer 2.5 generates world simulations conditioned on spatial controls (depth, segmentation, edges etc.) on top of Predict 2.5.","first":false,"discovered":"launch","source":"https://github.com/nvidia-cosmos/cosmos-transfer2.5"}],"entry":null,"notes":"Predict 2.5-2B released 2025-10-06 (per model card); needs ~32.5 GB VRAM. Consolidated into Cosmos 3 (June 2026) but still downloadable.","verified":"2026-09-29","body_md":"Sources: [Predict2.5-2B card](https://huggingface.co/nvidia/Cosmos-Predict2.5-2B), [Transfer2.5 GitHub](https://github.com/nvidia-cosmos/cosmos-transfer2.5).","page_url":"https://postcutoff.com/m/cosmos-predict-2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"sora-2","name":"Sora 2","org":"OpenAI","family":"Sora","released":"2025-10-06","status":"retired","type":"video-gen","modality_in":["text","image"],"modality_out":["video","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_second":0.1,"unit":"per second of video (720x1280 / 1280x720), before shutdown","source":"https://developers.openai.com/api/docs/models/sora-2"},"price_line":"$0.10 per second","access":[{"provider":"OpenAI API","model_id":"sora-2","endpoint":"https://api.openai.com/v1/videos","docs":"https://developers.openai.com/api/docs/models/sora-2"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"sora-2","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Synchronized audio","detail":"Generates video with audio from text or image prompts.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/sora-2"},{"name":"Shut down without replacement","detail":"Sora 2 models and Videos API shut down Sep 24 2026 with no one-to-one replacement.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/deprecations"}],"entry":null,"notes":"OpenAI API shut down 2026-09-24 (sora-2, sora-2-pro, snapshots sora-2-2025-10-06, sora-2-2025-12-08). Azure Foundry still listed sora-2 (preview) as of 2026-09-23. Release date = first API snapshot. Resellers followed: ElevenLabs removed Sora 2 and Sora 2 Pro from its Image & Video API on 2026-09-23 ('OpenAI is discontinuing the Sora API on September 24, 2026'), and the same changelog lists ByteDance retiring Seedance 1.5 Pro on 2026-11-11.","verified":"2026-09-29","body_md":"Retired on OpenAI API; kept for history. Azure Foundry listing may still work.\n\nSources:\n- https://developers.openai.com/api/docs/models/sora-2\n- https://developers.openai.com/api/docs/deprecations\n- https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure\n- ElevenLabs changelog 2026-09-23: https://elevenlabs.io/docs/changelog\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/sora-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"deepgram-flux","name":"Deepgram Flux (conversational STT, English + Multilingual)","org":"Deepgram","family":"Flux","released":"2025-10-02","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.0065,"unit":"per minute streaming, pay-as-you-go (flux-general-en); Flux Multilingual $0.0078/min; Growth plan $0.0057 / $0.0068","source":"https://deepgram.com/pricing"},"price_line":"$0.0065 per minute","access":[{"provider":"Deepgram API","model_id":"flux-general-en","endpoint":"wss://api.deepgram.com/v2/listen","docs":"https://developers.deepgram.com/docs/flux/quickstart"},{"provider":"Deepgram API","model_id":"flux-general-multi","endpoint":"wss://api.deepgram.com/v2/listen","docs":"https://developers.deepgram.com/docs/models-languages-overview"}],"capabilities":[{"name":"Conversational speech recognition with model-native turn-taking","detail":"Recognition model itself decides end-of-turn using acoustic + semantic cues (~260 ms end-of-turn detection), with EagerEndOfTurn events to start the LLM early; tunable eot_threshold, eager_eot_threshold, eot_timeout_ms.","first":true,"discovered":"launch","source":"https://deepgram.com/learn/introducing-flux-conversational-speech-recognition"},{"name":"Multilingual conversational STT with in-call code-switching","detail":"Flux Multilingual (GA 2026-04-29): English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian, Dutch with automatic language switching mid-conversation; turn detection under 400 ms. Billed by Deepgram as the world's first multilingual conversational speech recognition model.","first":true,"discovered":"later","source":"https://deepgram.com/learn/deepgram-launches-flux-multilingual-press-release"}],"entry":null,"notes":"'First' claims are Deepgram's own marketing (launched at VapiCon 2025-10-02 as 'world's first conversational speech recognition model'). Uses /v2/listen (not /v1). Mid-stream numeral toggle added 2026-09-25. Companion TTS: deepgram-flux-tts.","verified":"2026-09-29","body_md":"STT built for voice agents: transcription + end-of-turn detection in one model.\n\nSources: https://developers.deepgram.com/docs/flux/quickstart , https://deepgram.com/pricing , https://developers.deepgram.com/changelog","page_url":"https://postcutoff.com/m/deepgram-flux/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"gemini-2-5-flash-image","name":"Nano Banana (Gemini 2.5 Flash Image)","org":"Google DeepMind","family":"Gemini Image","released":"2025-10-02","status":"deprecated","type":"image-gen","modality_in":["text","image"],"modality_out":["image","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.3,"per_image":0.039,"unit":"input per 1M tokens; image output $0.039 per image","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.30 in per 1M tokens; $0.039 per image","access":[{"provider":"Gemini API","model_id":"gemini-2.5-flash-image","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-flash-image"},{"provider":"OpenRouter","model_id":"google/gemini-2.5-flash-image"}],"capabilities":[{"name":"Conversational image editing","detail":"The original 'Nano Banana': multi-turn natural-language image editing with character consistency, which made Gemini image editing go viral in 2025.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models"},{"name":"Multi-image fusion and targeted edits","detail":"Blend multiple images, keep characters consistent, and do prompt-based local edits (background blur, object removal, colorization); SynthID on all outputs.","first":false,"discovered":"launch","source":"https://developers.googleblog.com/en/introducing-gemini-2-5-flash-image/"}],"entry":null,"notes":"Stable GA 2025-10-02 (preview 2025-08-26); SHUTS DOWN 2026-10-02, replacement gemini-3.1-flash-image.","verified":"2026-09-29","body_md":"The first \"Nano Banana\" model. Migrate to `gemini-3.1-flash-image` now.\n\nSources: [models](https://ai.google.dev/gemini-api/docs/models), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/gemini-2-5-flash-image/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"neutts","name":"Neuphonic NeuTTS Air / NeuTTS Nano","org":"Neuphonic","family":"NeuTTS","released":"2025-10-02","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0 (NeuTTS Air)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"GitHub","url":"https://github.com/neuphonic/neutts"},{"provider":"Hugging Face","model_id":"neuphonic/neutts-nano-german","url":"https://huggingface.co/neuphonic/neutts-nano-german"}],"capabilities":[{"name":"On-device TTS with instant cloning","detail":"NeuTTS Air: 748M params (0.5B-class Qwen backbone + NeuCodec), real time from RTX 4090 down to Raspberry Pi, clones from ~3 s of audio, Perth watermark on every output; Nano: 229M total / 120M active for tighter edge devices.","first":false,"discovered":"launch","source":"https://www.marktechpost.com/2025/10/02/neuphonic-open-sources-neutts-air-a-748m-parameter-on-device-speech-language-model-with-instant-voice-cloning/"}],"entry":null,"notes":"Release date from MarkTechPost coverage (2025-10-02). Nano license and exact Air HF repo id (neuphonic/neutts-air) not verified today.","verified":null,"body_md":"Sources: https://github.com/neuphonic/neutts","page_url":"https://postcutoff.com/m/neutts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"hume-octave-2","name":"Hume Octave 2 (TTS)","org":"Hume AI","family":"Octave","released":"2025-10-01","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.15,"unit":"per 1K characters overage on Free/Starter/Creator plans; $0.12 Pro, $0.10 Scale, $0.05 Business (plans include monthly character quotas)","source":"https://www.hume.ai/pricing"},"price_line":"$0.15 per 1K characters","access":[{"provider":"Hume API","model_id":"version: 2","endpoint":"https://api.hume.ai/v0/tts","docs":"https://dev.hume.ai/docs/text-to-speech-tts/overview"},{"provider":"Web app","url":"https://platform.hume.ai"}],"capabilities":[{"name":"LLM-based emotionally intelligent TTS","detail":"Speech-language model that infers emotion and delivery from text; natural-language 'acting instructions' steer tone. Octave 2 at half the price of Octave 1, ~100 ms model latency (docs) / under 200 ms (launch blog).","first":false,"discovered":"launch","source":"https://www.hume.ai/blog/octave-2-launch"},{"name":"11 languages, instant cloning from 15 s","detail":"Arabic, English, French, German, Hindi, Italian, Japanese, Korean, Portuguese, Russian, Spanish; instant voice cloning from a ~15 s recording with accent prediction across languages; voice design from a text prompt (English only).","first":false,"discovered":"launch","source":"https://dev.hume.ai/docs/text-to-speech-tts/overview"},{"name":"Voice conversion and phoneme editing","detail":"Launch post describes voice conversion (swap speaker) and direct phoneme-level pronunciation editing as new capabilities for a speech-language model.","first":false,"discovered":"launch","source":"https://www.hume.ai/blog/octave-2-launch"}],"entry":null,"notes":"Select via `version: 2` in the TTS request body (`1` = Octave 1, English/Spanish, ~200 ms). Docs still label Octave 2 '(preview)' as of 2026-09-29. Auth header X-Hume-Api-Key. Max 5,000 chars per utterance, 1,000-char descriptions. Formats MP3/WAV/PCM. No Octave 3 announced on Hume's blog through Sept 2026. Speech-to-speech sibling: see hume-evi.","verified":"2026-09-29","body_md":"Emotionally expressive TTS from Hume AI (Octave = \"Omni-capable text and voice engine\").\n\n```bash\ncurl https://api.hume.ai/v0/tts -H \"X-Hume-Api-Key: $HUME_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"version\":2,\"utterances\":[{\"text\":\"I cannot believe it worked!\",\"description\":\"a delighted, breathless scientist\"}]}'\n```\n\nSources: https://dev.hume.ai/reference/text-to-speech-tts/synthesize-json , https://www.hume.ai/pricing , https://www.hume.ai/blog/octave-2-launch","page_url":"https://postcutoff.com/m/hume-octave-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"claude-sonnet-4-5","name":"Claude Sonnet 4.5","org":"Anthropic","family":"Claude 4","released":"2025-09-29","status":"deprecated","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":64000,"knowledge_cutoff":"2025-01","pricing":{"input":3,"output":15,"cache_read":0.3,"unit":"per 1M tokens (USD); Batch API 50% off","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$3 in, $15 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-sonnet-4-5-20250929","alias":"claude-sonnet-4-5","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/models/sonnet-4-5/overview"},{"provider":"AWS Bedrock (InvokeModel)","model_id":"anthropic.claude-sonnet-4-5-20250929-v1:0","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-amazon-bedrock-legacy"},{"provider":"Google Cloud Vertex AI","model_id":"claude-sonnet-4-5@20250929","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai"},{"provider":"Microsoft Foundry (Azure)","model_id":"claude-sonnet-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry"},{"provider":"Claude Platform on AWS","model_id":"claude-sonnet-4-5","docs":"https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws"},{"provider":"OpenRouter","model_id":"anthropic/claude-sonnet-4.5","url":"https://openrouter.ai/anthropic/claude-sonnet-4.5"},{"provider":"Web app","url":"https://claude.ai"}],"capabilities":[{"name":"30+ hour autonomous tasks","detail":"Anthropic reported it maintained focus for more than 30 hours on complex multi-step tasks.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-5"},{"name":"SOTA SWE-bench Verified at launch","detail":"77.2% on SWE-bench Verified; billed as 'the best coding model in the world' at release.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-5"},{"name":"Computer use lead","detail":"61.4% on OSWorld, up from 42.2% for Sonnet 4.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-sonnet-4-5"}],"entry":null,"notes":"Snapshot claude-sonnet-4-5-20250929 (alias claude-sonnet-4-5). Extended thinking only. Training data cutoff Jul 2025. Deprecated on 2026-09-30; retirement on the Claude API scheduled for 2026-11-30 (Anthropic recommends migrating to Claude Sonnet 5.5). Source: https://platform.claude.com/docs/en/release-notes/overview","verified":"2026-10-02","body_md":"Legacy Sonnet; migrate to `claude-sonnet-5-5`.\n\n```bash\ncurl https://api.anthropic.com/v1/messages \\\n -H \"x-api-key: $ANTHROPIC_API_KEY\" -H \"anthropic-version: 2023-06-01\" -H \"content-type: application/json\" \\\n -d '{\"model\":\"claude-sonnet-4-5\",\"max_tokens\":16000,\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources:\n- Model page: https://platform.claude.com/docs/en/models/sonnet-4-5/overview\n- Announcement: https://www.anthropic.com/news/claude-sonnet-4-5\n- Pricing: https://platform.claude.com/docs/en/about-claude/pricing\n- Models overview: https://platform.claude.com/docs/en/about-claude/models/overview\n- Deprecations: https://platform.claude.com/docs/en/about-claude/model-deprecations","page_url":"https://postcutoff.com/m/claude-sonnet-4-5/","events_after":881,"major_after":228,"historic_after":45,"missing_at_launch":43,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-01.md"}},{"id":"gemini-robotics-er-1","name":"Gemini Robotics-ER 1.5 / 1.6","org":"Google DeepMind","family":"Gemini Robotics","released":"2025-09-25","status":"retired","type":"robotics","modality_in":["text","image","video","audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":131072,"max_output":65536,"knowledge_cutoff":"2025-01","pricing":null,"price_line":"Price not published","access":[{"provider":"Gemini API (shut down)","model_id":"gemini-robotics-er-1.6-preview","docs":"https://ai.google.dev/gemini-api/docs/deprecations"},{"provider":"Gemini API (shut down)","model_id":"gemini-robotics-er-1.5-preview","docs":"https://ai.google.dev/gemini-api/docs/deprecations"}],"capabilities":[{"name":"Embodied reasoning in the public Gemini API","detail":"ER 1.5 (2025-09-25) exposed embodied reasoning (pointing, 2D boxes, trajectories, task planning, tool calls) available in the public Gemini API, while the VLA stayed partner-only.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/deprecations"},{"name":"Instrument reading (ER 1.6)","detail":"Reads pressure gauges, thermometers, sight glasses and digital readouts: 86% (93% with agentic vision) vs 23% for ER 1.5 and 67% for Gemini 3 Flash; built with Boston Dynamics and used by Spot for inspections.","first":false,"discovered":"launch","source":"https://deepmind.google/blog/gemini-robotics-er-1-6/"}],"entry":"2026-04-14-gemini-robotics-er-1-6","notes":"Retired. gemini-robotics-er-1.5-preview released 2025-09-25, shut down 2026-04-30 (replaced by 1.6). gemini-robotics-er-1.6-preview released 2026-04-14, shut down 2026-08-31 (replaced by gemini-robotics-er-2-preview). Token limits and Jan 2025 cutoff are those listed for ER 1.6 on the Gemini API model page.","verified":"2026-09-29","body_md":"Earlier embodied-reasoning previews; use [Gemini Robotics ER 2](gemini-robotics-er-2.md) instead.\n\nChangelog: 2026-09-29 added ER 1.6 instrument-reading capability and linked the new entry 2026-04-14-gemini-robotics-er-1-6.\n\nSources: [ER 1.6 blog](https://deepmind.google/blog/gemini-robotics-er-1-6/), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [ER 2 model page (lists 1.6 specs)](https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview).","page_url":"https://postcutoff.com/m/gemini-robotics-er-1/","events_after":881,"major_after":228,"historic_after":45,"missing_at_launch":42,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-01.md"}},{"id":"gemini-2-5-flash-native-audio","name":"Gemini 2.5 Flash Native Audio (Live, preview)","org":"Google DeepMind","family":"Gemini 2.5","released":"2025-09-23","status":"legacy","type":"audio/speech","modality_in":["text","audio","image","video"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.5,"output":2,"audio_input":3,"audio_output":12,"unit":"per 1M tokens (USD); audio/video in $3.00","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.50 in, $2 out per 1M tokens","access":[{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-2.5-flash-native-audio-preview-12-2025","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"Gemini Live API (WebSocket)","model_id":"gemini-2.5-flash-native-audio-preview-09-2025","docs":"https://ai.google.dev/gemini-api/docs/changelog"}],"capabilities":[{"name":"Native-audio reasoning in the Live API","detail":"Low-latency voice and video agents with native audio reasoning; 09-2025 snapshot improved function calling and speech cut-off handling, 12-2025 snapshot improved complex workflows.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/changelog"}],"entry":null,"notes":"Preview snapshots 2025-09-23 and 2025-12-12. No shutdown date announced; migrate to gemini-3.8-live. Older gemini-2.0-flash-live-001 and gemini-live-2.5-flash-preview were shut down 2025-12-09.","verified":"2026-09-29","body_md":"Sources: [models](https://ai.google.dev/gemini-api/docs/models), [changelog](https://ai.google.dev/gemini-api/docs/changelog), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [pricing](https://ai.google.dev/gemini-api/docs/pricing).","page_url":"https://postcutoff.com/m/gemini-2-5-flash-native-audio/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":839,"you_url":null,"briefings":null},{"id":"indextts-2","name":"IndexTTS-2 / IndexTTS-2.5 (bilibili)","org":"bilibili (Index Team)","family":"IndexTTS","released":"2025-09-08","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"bilibili Model Use License Agreement (commercial use by request)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"IndexTeam/IndexTTS-2","url":"https://huggingface.co/IndexTeam/IndexTTS-2"},{"provider":"GitHub","url":"https://github.com/index-tts/index-tts"}],"capabilities":[{"name":"Precise duration control in an autoregressive TTS","detail":"Lets users specify the exact number of speech tokens (useful for dubbing/lip-sync) while keeping AR naturalness, and disentangles speaker timbre from emotion (emotion from a separate reference audio or text).","first":false,"discovered":"launch","source":"https://huggingface.co/IndexTeam/IndexTTS-2"},{"name":"IndexTTS-2.5 multilingual","detail":"2026-08-10 release adds Japanese, Spanish and Arabic to Chinese/English; speed 0.5-2x, Pinyin/CMU/Kana pronunciation control, RTF ~0.2 on RTX 4090.","first":false,"discovered":"later","source":"https://github.com/index-tts/index-tts"}],"entry":null,"notes":"The IndexTTS2 paper (arXiv June 2025) presents duration control as novel for AR TTS; 'first' not independently verified, so not flagged. Weights released 2025-09-08. Commercial use: contact indexspeech@bilibili.com.","verified":"2026-09-29","body_md":"Sources: https://github.com/index-tts/index-tts , https://huggingface.co/IndexTeam/IndexTTS-2","page_url":"https://postcutoff.com/m/indextts-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":845,"you_url":null,"briefings":null},{"id":"elevenlabs-sound-effects-v2","name":"Eleven Sound Effects v2","org":"ElevenLabs","family":"Eleven Sound Effects","released":"2025-09","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.12,"unit":"per minute of generated audio (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.12 per minute","access":[{"provider":"ElevenLabs API","model_id":"eleven_text_to_sound_v2","endpoint":"https://api.elevenlabs.io/v1/sound-generation","docs":"https://elevenlabs.io/docs/overview/capabilities/sound-effects"},{"provider":"fal","model_id":"fal-ai/elevenlabs/sound-effects/v2","url":"https://fal.ai/models/fal-ai/elevenlabs/sound-effects/v2"},{"provider":"Web app","url":"https://elevenlabs.io/sound-effects"}],"capabilities":[{"name":"Seamless looping SFX, 48 kHz","detail":"Text-to-sound effects up to 30 s per generation (0.1-30 s selectable), seamless looping for longer ambiences, prompt-influence control; MP3, WAV 48 kHz for non-looping.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/overview/capabilities/sound-effects"}],"entry":null,"notes":"Release month (Sept 2025) is from third-party sources, not an official post. App pricing: 40 credits/second when duration is set.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/docs/overview/capabilities/sound-effects","page_url":"https://postcutoff.com/m/elevenlabs-sound-effects-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"rdt2","name":"RDT2 (and RDT-1B)","org":"Tsinghua University (TSAIL, thu-ml)","family":"Robotics Diffusion Transformer","released":"2025-09","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"robotics-diffusion-transformer/RDT2-VQ","url":"https://huggingface.co/robotics-diffusion-transformer/RDT2-VQ"},{"provider":"Hugging Face (RDT-1B, MIT)","model_id":"robotics-diffusion-transformer/rdt-1b","url":"https://huggingface.co/robotics-diffusion-transformer/rdt-1b"},{"provider":"GitHub","url":"https://github.com/thu-ml/RDT2","docs":"https://rdt-robotics.github.io/rdt2/"}],"capabilities":[{"name":"Zero-shot deployment on unseen embodiments","detail":"RDT2 (8B, Qwen2.5-VL-7B based, residual-VQ action tokens; RDT2-FM flow-matching variant) trained on 10k+ h of UMI-gripper human manipulation from 100+ scenes; authors call it possibly the first foundation model to deploy zero-shot on unseen embodiments (UR5e, Franka FR3) for simple open-vocabulary tasks.","first":true,"discovered":"launch","source":"https://huggingface.co/robotics-diffusion-transformer/RDT2-VQ"},{"name":"Large diffusion foundation model for bimanual manipulation (RDT-1B)","detail":"RDT-1B (Oct 2024, 1.2B) was billed as the largest diffusion-based foundation model for bimanual manipulation, pretrained on 46 datasets (1M+ episodes) and fine-tuned on a 6K+ episode ALOHA dataset.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2410.07864"}],"entry":null,"notes":"'First' claim is the authors' own hedged wording. HF RDT2-VQ repo created 2025-09-22.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/rdt2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":834,"you_url":null,"briefings":null},{"id":"agility-digit-motor-cortex","name":"Agility Digit whole-body control foundation model (\"motor cortex\")","org":"Agility Robotics","family":"Agility Arc / Digit AI","released":"2025-08-28","status":"current","type":"robotics","modality_in":["text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Onboard Agility Digit (commercial humanoid, via Agility)","url":"https://www.agilityrobotics.com/content/agility-and-ai","docs":"https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model"}],"capabilities":[{"name":"Tiny sim-trained whole-body controller","detail":"An LSTM with fewer than 1M parameters, trained with RL in NVIDIA Isaac Sim for decades of simulated time in 3-4 days, transferring zero-shot to Digit for balance, walking, arm placement and carrying heavy objects while staying stable.","first":false,"discovered":"launch","source":"https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model"},{"name":"Layered stack with LLM on top","detail":"Higher layers (open-vocabulary detectors, state-machine planners, an LLM such as a Gemini research preview) send targets to the motor cortex; dexterous skills are learned on top of it.","first":false,"discovered":"launch","source":"https://www.agilityrobotics.com/content/training-a-whole-body-control-foundation-model"}],"entry":null,"notes":"Agility has not published a large VLA of its own; this is its disclosed foundation-model layer. Digit is in paid deployments (e.g. GXO); Agility opened a Fremont \"Physical AI\" facility in July 2026 (https://www.nasdaq.com/press-release/agility-opens-new-fremont-facility-accelerate-physical-ai-development-2026-07-16). Not developer-accessible.","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/agility-digit-motor-cortex/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":846,"you_url":null,"briefings":null},{"id":"gpt-realtime","name":"GPT-Realtime and GPT-Realtime mini","org":"OpenAI","family":"GPT Realtime","released":"2025-08-28","status":"deprecated","type":"audio/speech","modality_in":["text","audio","image"],"modality_out":["text","audio"],"open_weights":false,"model_license":"proprietary","context_window":32000,"max_output":4096,"knowledge_cutoff":"2023-10","pricing":{"text_input":4,"text_output":16,"audio_input":32,"audio_output":64,"cached_input":0.4,"unit":"per 1M tokens (USD) for gpt-realtime; gpt-realtime-mini audio $10 in / $20 out, text $0.60 / $2.40","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$4 text in, $16 text out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-realtime","endpoint":"wss://api.openai.com/v1/realtime","docs":"https://developers.openai.com/api/docs/models/gpt-realtime"},{"provider":"OpenAI API","model_id":"gpt-realtime-mini","endpoint":"wss://api.openai.com/v1/realtime","docs":"https://developers.openai.com/api/docs/models/gpt-realtime-mini"}],"capabilities":[{"name":"First GA OpenAI realtime speech-to-speech model","detail":"Shipped with Realtime API general availability (2025-08-28); speaks over WebRTC, WebSocket or SIP phone calls.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-realtime"}],"entry":null,"notes":"Deprecated 2026-07-20, shutdown 2027-01-20; replacements gpt-realtime-2.1 and gpt-realtime-2.1-mini. gpt-realtime-mini released 2025-10-06; its alias moved to the 2025-12-15 snapshot on 2026-01-13. Earlier gpt-4o-realtime-preview models were shut down 2026-05-12.","verified":"2026-09-29","body_md":"Sources:\n- https://developers.openai.com/api/docs/models/gpt-realtime\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/changelog\n- https://developers.openai.com/api/docs/deprecations","page_url":"https://postcutoff.com/m/gpt-realtime/","events_after":930,"major_after":256,"historic_after":53,"missing_at_launch":84,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2023-10-xs.md","s":"https://postcutoff.com/briefings/cutoff-2023-10-s.md","m":"https://postcutoff.com/briefings/cutoff-2023-10-m.md","full":"https://postcutoff.com/briefings/cutoff-2023-10.md"}},{"id":"vibevoice","name":"VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS)","org":"Microsoft","family":"VibeVoice","released":"2025-08-25","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text","audio"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"microsoft/VibeVoice-ASR","url":"https://huggingface.co/microsoft/VibeVoice-ASR"},{"provider":"Hugging Face (Transformers format)","model_id":"microsoft/VibeVoice-ASR-HF","url":"https://huggingface.co/microsoft/VibeVoice-ASR-HF"},{"provider":"Hugging Face","model_id":"microsoft/VibeVoice-ASR-BitNet","url":"https://huggingface.co/microsoft/VibeVoice-ASR-BitNet"},{"provider":"Hugging Face","model_id":"microsoft/VibeVoice-Realtime-0.5B","url":"https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B"},{"provider":"Hugging Face (streaming ASR, 7B repo; 9B params incl. decoder)","model_id":"microsoft/VibeVoice-ASR-Streaming-7B","url":"https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B"},{"provider":"Hugging Face (streaming ASR, small)","model_id":"microsoft/VibeVoice-ASR-Streaming-1.5B","url":"https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-1.5B"},{"provider":"GitHub","url":"https://github.com/microsoft/VibeVoice"}],"capabilities":[{"name":"60-minute single-pass ASR with diarization","detail":"VibeVoice-ASR (~9B params incl. Qwen2-based decoder) transcribes up to 60 min in one pass with who/when/what structured output, hotwords and 50+ languages with code-switching.","first":false,"discovered":"launch","source":"https://huggingface.co/microsoft/VibeVoice-ASR"},{"name":"CPU-only realtime ASR","detail":"VibeVoice-ASR-BitNet (2026-07-23) compresses the model 4.62 GB -> 1.58 GB and runs faster than real time on 3 CPU threads (1.6-2.3x faster than Whisper.cpp).","first":false,"discovered":"later","source":"https://huggingface.co/microsoft/VibeVoice-ASR-BitNet"},{"name":"Streaming speaker-attributed ASR","detail":"VibeVoice-ASR-Streaming (7B and 1.5B repos, uploaded 2026-09-02) transcribes live audio with speaker attribution (who said what) and custom hotwords in 10 languages (zh, en, fr, de, it, ja, ko, pt, ru, es); MIT license.","first":false,"discovered":"later","source":"https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B"},{"name":"Long-form multi-speaker TTS (withdrawn)","detail":"Original VibeVoice-TTS (1.5B/7B) generated up to 90 min with 4 speakers; Microsoft removed the TTS code on 2025-09-05 over responsible-AI misuse concerns.","first":false,"discovered":"later","source":"https://github.com/microsoft/VibeVoice"}],"entry":null,"notes":"Open-source voice research family from Microsoft (MIT). Timeline: TTS 2025-08-25 (code pulled 2025-09-05), Realtime-0.5B streaming TTS (~300 ms first audio) 2025-12-03, ASR 2026-01-21, Transformers integration 2026-03, Foundry Labs 2026-03-12, ASR-BitNet 2026-07-23, ASR-Streaming (10 languages, hotwords, speaker attribution) announced 2026-09-03; HF repos microsoft/VibeVoice-ASR-Streaming-7B and -1.5B created 2026-09-02 (verified 2026-09-29). Monthly downloads to 2026-09-29: VibeVoice-ASR ~734k, VibeVoice-1.5B ~717k. Separate from Microsoft's proprietary MAI-Voice/MAI-Transcribe.","verified":"2026-09-29","body_md":"Open Microsoft speech models for long-form ASR and lightweight streaming TTS.\n\nSee the Hugging Face Transformers docs (model_doc/vibevoice_asr) and the GitHub repo for inference code.\n\nSources: https://github.com/microsoft/VibeVoice · https://huggingface.co/microsoft/VibeVoice-ASR · https://huggingface.co/docs/transformers/model_doc/vibevoice_asr","page_url":"https://postcutoff.com/m/vibevoice/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":847,"you_url":null,"briefings":null},{"id":"atlas-large-behavior-model","name":"Atlas Large Behavior Model (Boston Dynamics + TRI LBM)","org":"Boston Dynamics","family":"Large Behavior Models (TRI)","released":"2025-08-20","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Not available (internal research policy for Atlas)","url":"https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/"},{"provider":"TRI LBM Eval (open simulation benchmark, not the model)","url":"https://github.com/ToyotaResearchInstitute/lbm_eval"}],"capabilities":[{"name":"One language-conditioned policy for whole-body loco-manipulation","detail":"A single end-to-end policy maps images, proprioception and language to actions for the full 50-DoF Atlas at 30 Hz, combining stepping, crouching and center-of-mass shifts with dexterous manipulation in long-horizon tasks, replacing separate walking and manipulation controllers.","first":false,"discovered":"launch","source":"https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/"},{"name":"Diffusion Transformer with flow matching","detail":"450M-parameter Diffusion Transformer trained with a flow-matching objective, predicting 48-step action chunks (1.6 s); trained on Atlas teleop data, the Atlas Manipulation Test Stand, TRI's Ramen dataset and simulation co-training.","first":false,"discovered":"launch","source":"https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/"},{"name":"Inference-time speed-up","detail":"Policies can run 1.5-2x faster than the human demos at inference time without retraining by rescaling action timing.","first":false,"discovered":"launch","source":"https://bostondynamics.com/blog/large-behavior-models-atlas-find-new-footing/"},{"name":"Pretraining cuts task data by up to 80%","detail":"TRI's LBM study (~1,700 h of robot data, 1,800 real and 47,000+ sim rollouts) found pretrained LBMs need up to 80% less task-specific data.","first":false,"discovered":"launch","source":"https://toyotaresearchinstitute.github.io/lbm1/"}],"entry":"2026-01-05-boston-dynamics-atlas-production","notes":"Research collaboration announced Aug 2025 (Toyota release: https://newsroom.toyota.eu/ai-powered-robot-by-boston-dynamics-and-toyota-research-institute-takes-a-key-step-towards-general-purpose-humanoids/). The production electric Atlas (CES 2026) also integrates Google DeepMind foundation models (Gemini Robotics); Hyundai trains Atlas on parts logistics at its Georgia RMAC (2026-09-22). No public weights or API for the Atlas LBM. Exact announcement day (2025-08-20) is from press coverage dated 2025-08-20/21.","verified":null,"body_md":null,"page_url":"https://postcutoff.com/m/atlas-large-behavior-model/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":847,"you_url":null,"briefings":null},{"id":"nvidia-parakeet-canary","name":"NVIDIA Parakeet / Canary / Nemotron Speech ASR (open)","org":"NVIDIA","family":"NeMo ASR","released":"2025-08-14","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"cc-by-4.0 (Parakeet TDT v3, Canary-Qwen); NVIDIA Open Model License / OpenMDW (2026 models)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/parakeet-tdt-0.6b-v3","url":"https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3"},{"provider":"Hugging Face","model_id":"nvidia/canary-qwen-2.5b","url":"https://huggingface.co/nvidia/canary-qwen-2.5b"},{"provider":"Hugging Face","model_id":"nvidia/parakeet-unified-en-0.6b","url":"https://huggingface.co/nvidia/parakeet-unified-en-0.6b"},{"provider":"Hugging Face","model_id":"nvidia/nemotron-speech-streaming-en-0.6b","url":"https://huggingface.co/nvidia/nemotron-speech-streaming-en-0.6b"},{"provider":"Hugging Face","model_id":"nvidia/nemotron-3.5-asr-streaming-0.6b","url":"https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b"},{"provider":"NVIDIA NIM / build.nvidia.com","url":"https://build.nvidia.com"}],"capabilities":[{"name":"Parakeet TDT 0.6B v3: 25 European languages, very high throughput","detail":"600M FastConformer-TDT with auto language ID, punctuation, word timestamps, up to 24 min (3 h with local attention); 6.34% avg WER on Open ASR Leaderboard; trained on the Granary dataset.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3"},{"name":"Canary-Qwen-2.5B speech-augmented LLM","detail":"FastConformer encoder + Qwen LLM (SALM); 5.63% mean WER topped the HF Open ASR Leaderboard at release (2025-07-17); can summarize/answer questions about the transcript. English, max 40 s clips.","first":false,"discovered":"launch","source":"https://huggingface.co/nvidia/canary-qwen-2.5b"},{"name":"Cache-aware streaming ASR, 80-1120 ms chunks","detail":"Nemotron Speech Streaming EN 0.6B (Jan/Mar 2026) and Nemotron 3.5 ASR Streaming 0.6B (June 2026, 40 language-locales) switch latency at inference without retraining; Parakeet-unified-en-0.6B (2026-04-07) does both offline (5.91% WER) and streaming down to 160 ms.","first":false,"discovered":"later","source":"https://huggingface.co/nvidia/nemotron-3.5-asr-streaming-0.6b"}],"entry":null,"notes":"One file for NVIDIA's open ASR family. Also canary-1b-v2 (European ASR + translation) and parakeet-tdt-0.6b-v2 (English, NIM). Nemotron 3.5 ASR HF card shows a garbled date; June 2026 per NVIDIA/press. NVIDIA's open TTS: magpie_tts_multilingual_357m. Full-duplex model: see nemotron-voicechat.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/collections/nvidia/nemotron-speech , https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 , https://huggingface.co/nvidia/canary-qwen-2.5b","page_url":"https://postcutoff.com/m/nvidia-parakeet-canary/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":848,"you_url":null,"briefings":null},{"id":"claude-opus-4-1","name":"Claude Opus 4.1","org":"Anthropic","family":"Claude 4","released":"2025-08-05","status":"retired","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":15,"output":75,"unit":"per 1M tokens (USD), Bedrock/Google Cloud may differ","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$15 in, $75 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-opus-4-1-20250805","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/about-claude/model-deprecations"},{"provider":"OpenRouter","model_id":"anthropic/claude-opus-4.1","url":"https://openrouter.ai/anthropic/claude-opus-4.1"}],"capabilities":[{"name":"SOTA SWE-bench Verified (Aug 2025)","detail":"74.5% on SWE-bench Verified at launch.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-1"},{"name":"Precise multi-file refactoring","detail":"GitHub and Rakuten highlighted multi-file refactoring and pinpoint fixes without unnecessary changes.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-opus-4-1"}],"entry":null,"notes":"Retired on the Claude API 2026-08-05 (replacement claude-opus-4-8 / claude-opus-5-5); still available on Amazon Bedrock and Google Cloud per Anthropic pricing page. Cloud ids not re-verified today.","verified":"2026-09-29","body_md":"Retired on Anthropic-operated platforms; requests to `claude-opus-4-1-20250805` on the Claude API fail. Listed for reference because it may still be reachable via Bedrock, Google Cloud or OpenRouter.\n\nSources: https://platform.claude.com/docs/en/about-claude/model-deprecations , https://platform.claude.com/docs/en/about-claude/pricing","page_url":"https://postcutoff.com/m/claude-opus-4-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":851,"you_url":null,"briefings":null},{"id":"gpt-oss-120b","name":"gpt-oss-120b","org":"OpenAI","family":"gpt-oss","released":"2025-08-05","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":131072,"max_output":131072,"knowledge_cutoff":"2024-06","pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/openai/gpt-oss-120b"},{"provider":"OpenAI API (docs)","model_id":"gpt-oss-120b","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-oss-120b"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-oss-120b","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-oss-120b","url":"https://openrouter.ai/openai/gpt-oss-120b"}],"capabilities":[{"name":"Single-GPU open MoE","detail":"117B total / 5.1B active MoE with MXFP4 weights; runs on one 80GB H100/MI300X.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/gpt-oss-120b"},{"name":"Open reasoning with full CoT","detail":"Configurable low/medium/high reasoning with full chain-of-thought access, harmony format.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/gpt-oss-120b"}],"entry":null,"notes":"Open weights (Apache 2.0). OpenRouter from ~$0.04/$0.17 per 1M (provider-dependent). No first-party OpenAI pricing listed.","verified":"2026-09-29","body_md":"Self-hosted reasoning/agents via vLLM, Transformers, Ollama, LM Studio.\n\n```bash\nvllm serve openai/gpt-oss-120b\n```\n\nSources:\n- https://huggingface.co/openai/gpt-oss-120b\n- https://developers.openai.com/api/docs/models/gpt-oss-120b\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-oss-120b/","events_after":904,"major_after":242,"historic_after":53,"missing_at_launch":51,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-oss-120b/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-06.md"}},{"id":"gpt-oss-20b","name":"gpt-oss-20b","org":"OpenAI","family":"gpt-oss","released":"2025-08-05","status":"current","type":"reasoning-llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":131072,"max_output":131072,"knowledge_cutoff":"2024-06","pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/openai/gpt-oss-20b"},{"provider":"OpenAI API (docs)","model_id":"gpt-oss-20b","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/gpt-oss-20b"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-oss-20b","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-oss-20b","url":"https://openrouter.ai/openai/gpt-oss-20b"}],"capabilities":[{"name":"Laptop-class open reasoning","detail":"21B total / 3.6B active MoE in MXFP4; runs in ~16GB memory.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/gpt-oss-20b"},{"name":"Fine-tunable on consumer hardware","detail":"Apache 2.0 weights, fine-tunable locally; function calling and structured outputs.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/gpt-oss-20b"}],"entry":null,"notes":"Open weights (Apache 2.0). Azure lists it as Preview. Safety-classifier variant openai/gpt-oss-safeguard-20b also on OpenRouter.","verified":"2026-09-29","body_md":"Local/on-device reasoning.\n\n```bash\nollama run gpt-oss:20b\n```\n\nSources:\n- https://huggingface.co/openai/gpt-oss-20b\n- https://developers.openai.com/api/docs/models/gpt-oss-20b\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-oss-20b/","events_after":904,"major_after":242,"historic_after":53,"missing_at_launch":51,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-oss-20b/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-06.md"}},{"id":"genie-3","name":"Genie 3","org":"Google DeepMind","family":"Genie","released":"2025-08","status":"preview","type":"world-model","modality_in":["text","image"],"modality_out":["video"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Project Genie (Google Labs)","url":"https://labs.google/projectgenie","docs":"https://deepmind.google/models/genie/"}],"capabilities":[{"name":"Real-time interactive world generation","detail":"Generates navigable, photorealistic 720p worlds at 20-24 fps from text/image prompts.","first":true,"discovered":"launch","source":"https://deepmind.google/models/genie/"},{"name":"World memory / consistency","detail":"Regions stay consistent when revisited; multi-minute visual consistency.","first":false,"discovered":"launch","source":"https://deepmind.google/models/genie/"},{"name":"Promptable world events","detail":"Change weather or introduce objects/characters mid-exploration via text.","first":false,"discovered":"launch","source":"https://deepmind.google/models/genie/"},{"name":"Consumer world sketching and remixing","detail":"Project Genie (2026-01-29) lets users sketch, explore and remix worlds (60 s sessions), combining Genie 3 with Nano Banana Pro and Gemini.","first":false,"discovered":"later","source":"https://blog.google/innovation-and-ai/models-and-research/google-deepmind/project-genie/"}],"entry":null,"notes":"No public API. Research preview announced Aug 2025; consumer access via Project Genie only for Google AI Ultra subscribers, US, 18+ (not Business accounts). Exact announcement day not re-verified.","verified":null,"body_md":"DeepMind's general-purpose world model for interactive environments and agent training. Not callable via API; try it at labs.google/projectgenie (AI Ultra).\n\nSources: [Genie 3 page](https://deepmind.google/models/genie/), [Project Genie blog](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/project-genie/).","page_url":"https://postcutoff.com/m/genie-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":846,"you_url":null,"briefings":null},{"id":"codestral-2508","name":"Codestral 25.08","org":"Mistral AI","family":"Codestral","released":"2025-07-30","status":"current","type":"code","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"mistral-premier","context_window":128000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.3,"output":0.9,"unit":"per 1M tokens (USD)","source":"https://docs.mistral.ai/models/model-cards/codestral-25-08"},"price_line":"$0.30 in, $0.90 out per 1M tokens","access":[{"provider":"Mistral API (FIM)","model_id":"codestral-2508","endpoint":"https://api.mistral.ai/v1/fim/completions","docs":"https://docs.mistral.ai/models/model-cards/codestral-25-08"},{"provider":"Mistral API (chat)","model_id":"codestral-latest","endpoint":"https://api.mistral.ai/v1/chat/completions"},{"provider":"OpenRouter","model_id":"mistralai/codestral-2508","url":"https://openrouter.ai/mistralai/codestral-2508"}],"capabilities":[{"name":"Low-latency fill-in-the-middle","detail":"Specialized for high-frequency FIM/autocomplete with a dedicated FIM endpoint, plus predicted outputs.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/codestral-25-08"},{"name":"Predicted outputs and prefix mode","detail":"Supports predicted outputs (fast edits of known code) and assistant prefix, plus function calling and structured outputs.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/model-cards/codestral-25-08"}],"entry":null,"notes":"Alias codestral-latest. Mistral's current code-completion model (Premier). OpenRouter lists 256K context; Mistral card says 128K.","verified":"2026-09-29","body_md":"IDE autocomplete / FIM and fast code generation.\n\n```bash\ncurl https://api.mistral.ai/v1/fim/completions -H \"Authorization: Bearer $MISTRAL_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"codestral-2508\",\"prompt\":\"def fib(n):\",\"suffix\":\"\\nprint(fib(10))\"}'\n```\n\nSources: https://docs.mistral.ai/models/model-cards/codestral-25-08","page_url":"https://postcutoff.com/m/codestral-2508/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":853,"you_url":null,"briefings":null},{"id":"gemini-2-5-flash-lite","name":"Gemini 2.5 Flash-Lite","org":"Google DeepMind","family":"Gemini 2.5","released":"2025-07-22","status":"legacy","type":"llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.1,"output":0.4,"unit":"per 1M tokens (Standard; text/image/video input; audio input $0.30)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.10 in, $0.40 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-2.5-flash-lite","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"OpenRouter","model_id":"google/gemini-2.5-flash-lite"}],"capabilities":[{"name":"Cheapest Gemini text tier","detail":"Still the lowest per-token Gemini text price ($0.10 / $0.40) with 1M context.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/pricing"},{"name":"Thinking off by default","detail":"Lowest latency/cost in the 2.5 family, thinking disabled by default, yet supports grounding, code execution, URL context and function calling.","first":false,"discovered":"launch","source":"https://developers.googleblog.com/en/gemini-2-5-thinking-model-updates/"}],"entry":null,"notes":"GA 2025-07-22; no shutdown date, access limited to historical users; replacement 3.5 Flash-Lite. Max output not re-verified.","verified":"2026-09-29","body_md":"Budget 2025 model; legacy.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [models](https://ai.google.dev/gemini-api/docs/models), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/gemini-2-5-flash-lite/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":857,"you_url":null,"briefings":null},{"id":"kyutai-tts-stt","name":"Kyutai TTS 1.6B / Kyutai STT + Unmute","org":"Kyutai","family":"Delayed Streams Modeling","released":"2025-07-03","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio","text"],"open_weights":true,"model_license":"cc-by-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face (TTS)","model_id":"kyutai/tts-1.6b-en_fr","url":"https://huggingface.co/kyutai/tts-1.6b-en_fr"},{"provider":"Hugging Face (STT)","model_id":"kyutai/stt-2.6b-en","url":"https://huggingface.co/kyutai/stt-2.6b-en"},{"provider":"Hugging Face (STT)","model_id":"kyutai/stt-1b-en_fr","url":"https://huggingface.co/kyutai/stt-1b-en_fr"},{"provider":"GitHub (Unmute)","url":"https://github.com/kyutai-labs/unmute"}],"capabilities":[{"name":"Text-streaming TTS","detail":"Delayed-streams architecture (~1.8B params incl. 600M depth transformer) starts speaking before the full text is available, English + French; voices only via pre-computed embeddings (no raw cloning, by design).","first":false,"discovered":"launch","source":"https://huggingface.co/kyutai/tts-1.6b-en_fr"},{"name":"Streaming STT with semantic VAD","detail":"stt-2.6b-en (English, 2.5 s delay) and stt-1b-en_fr (0.5 s delay) transcribe as audio arrives; used in Unmute, which wraps any text LLM with real-time STT+TTS.","first":false,"discovered":"launch","source":"https://huggingface.co/kyutai/stt-2.6b-en"}],"entry":null,"notes":"STT open-sourced 2025-06-19, TTS + Unmute open-sourced 2025-07-03 (Kyutai blog). Weights CC-BY-4.0. For CPU TTS see kyutai-pocket-tts.","verified":"2026-09-29","body_md":"Sources: https://kyutai.org/blog/ , https://huggingface.co/kyutai/tts-1.6b-en_fr , https://huggingface.co/kyutai/stt-2.6b-en","page_url":"https://postcutoff.com/m/kyutai-tts-stt/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":860,"you_url":null,"briefings":null},{"id":"voxtral-small","name":"Voxtral Small","org":"Mistral AI","family":"Voxtral","released":"2025-07","status":"current","type":"multimodal","modality_in":["audio","text"],"modality_out":["text"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Mistral API","model_id":"voxtral-small-2507","endpoint":"https://api.mistral.ai/v1/chat/completions","docs":"https://docs.mistral.ai/models/overview"},{"provider":"Hugging Face","model_id":"mistralai/Voxtral-Small-24B-2507","url":"https://huggingface.co/mistralai/Voxtral-Small-24B-2507"}],"capabilities":[{"name":"Audio-understanding chat model","detail":"Mistral's first model with audio input for instruct use (Q&A, summarization, function calling from voice) on top of transcription.","first":false,"discovered":"launch","source":"https://docs.mistral.ai/models/overview"}],"entry":null,"notes":"Still listed as active (v25.07) on Mistral's models overview on 2026-09-29; its small siblings voxtral-mini-2507 and Voxtral Mini Transcribe 25.07 were retired 2026-05-31. Pricing and exact release day not re-verified (July 2025 launch).","verified":"2026-09-29","body_md":"Open 24B audio-in LLM from Mistral (July 2025).\n\nSources: https://docs.mistral.ai/models/overview · https://huggingface.co/mistralai/Voxtral-Small-24B-2507","page_url":"https://postcutoff.com/m/voxtral-small/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":853,"you_url":"https://postcutoff.com/you/voxtral-small/","briefings":null},{"id":"imagen-4","name":"Imagen 4","org":"Google DeepMind","family":"Imagen","released":"2025-06-24","status":"retired","type":"image-gen","modality_in":["text"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Gemini API (shut down)","model_id":"imagen-4.0-generate-001","docs":"https://ai.google.dev/gemini-api/docs/deprecations"}],"capabilities":[{"name":"Dedicated text-to-image diffusion model","detail":"Google's last standalone Imagen generation; superseded by Gemini-native image models (Nano Banana 2).","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/deprecations"},{"name":"Retired in favor of Gemini-native imaging","detail":"Deprecation table names gemini-3.1-flash-image as replacement, marking the shift from standalone diffusion models to Gemini image models.","first":false,"discovered":"later","source":"https://ai.google.dev/gemini-api/docs/deprecations"}],"entry":null,"notes":"Imagen 4.0 variants released 2025-06-24, shut down in the Gemini API 2026-08-17; replacement gemini-3.1-flash-image. Other variant ids (fast/ultra) and Vertex status not verified; pricing not verified (retired).","verified":"2026-09-29","body_md":"Retired. Use `gemini-3.1-flash-image` (Nano Banana 2) or `gemini-3-pro-image` instead.\n\nSources: [deprecations](https://ai.google.dev/gemini-api/docs/deprecations), [models](https://ai.google.dev/gemini-api/docs/models).","page_url":"https://postcutoff.com/m/imagen-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":861,"you_url":null,"briefings":null},{"id":"gemini-2-5-flash","name":"Gemini 2.5 Flash","org":"Google DeepMind","family":"Gemini 2.5","released":"2025-06-17","status":"legacy","type":"reasoning-llm","modality_in":["text","image","audio","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":"2025-01","pricing":{"input":0.3,"output":2.5,"unit":"per 1M tokens (Standard; text/image/video input; audio input $1.00)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.30 in, $2.50 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-2.5-flash","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"gemini-2.5-flash","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-flash"},{"provider":"OpenRouter","model_id":"google/gemini-2.5-flash"}],"capabilities":[{"name":"Hybrid reasoning with thinking budget","detail":"Thinking can be controlled per request; 1M-token multimodal context at low price.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash"},{"name":"First fully hybrid reasoning model (Google)","detail":"Google's first model where thinking can be switched on/off, with a 0-24,576 token thinking budget.","first":false,"discovered":"launch","source":"https://developers.googleblog.com/en/start-building-with-gemini-25-flash/"}],"entry":null,"notes":"Preview 2025-04-17, GA 2025-06-17. No shutdown date, but access limited to prior users; replacement 3.5 Flash-Lite or 3.8 Flash.","verified":"2026-09-29","body_md":"2025 workhorse model; legacy. Use `gemini-3.8-flash` or `gemini-3.5-flash-lite` for new work.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/gemini-2-5-flash/","events_after":881,"major_after":228,"historic_after":45,"missing_at_launch":19,"events_since_release":null,"you_url":"https://postcutoff.com/you/gemini-2-5-flash/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-01.md"}},{"id":"gemini-2-5-pro","name":"Gemini 2.5 Pro","org":"Google DeepMind","family":"Gemini 2.5","released":"2025-06-17","status":"legacy","type":"reasoning-llm","modality_in":["text","image","audio","video","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1048576,"max_output":65536,"knowledge_cutoff":"2025-01","pricing":{"input":1.25,"output":10,"unit":"per 1M tokens (Standard, prompts <=200k; >200k: $2.50 in / $15.00 out)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$1.25 in, $10 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-2.5-pro","endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-pro:generateContent","docs":"https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro"},{"provider":"Google Cloud Gemini Enterprise Agent Platform (Vertex AI)","model_id":"gemini-2.5-pro","docs":"https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-pro"},{"provider":"OpenRouter","model_id":"google/gemini-2.5-pro"}],"capabilities":[{"name":"Thinking model with 1M context","detail":"Built-in thinking plus 1,048,576-token multimodal input and 65K output; Search/Maps grounding, code execution, URL context.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro"},{"name":"Debuted","detail":"First Gemini 2.5 'thinking model'; the March 2025 experimental release topped LMArena by a significant margin and led coding/math/science benchmarks.","first":false,"discovered":"launch","source":"https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-model-thinking-updates-march-2025/"},{"name":"Coding-agent backbone","detail":"Steepest demand growth of any Google model; powered tools such as Cursor and GitHub Copilot at GA.","first":false,"discovered":"later","source":"https://developers.googleblog.com/en/gemini-2-5-thinking-model-updates/"}],"entry":null,"notes":"First released as experimental 2025-03; GA 2025-06-17 (stable id; earlier preview ids e.g. gemini-2.5-pro-preview-*). No shutdown date, but Gemini API access is limited to projects that used it before; Google recommends 3.5 Flash-Lite or 3.8 Flash for new projects.","verified":"2026-09-29","body_md":"Former flagship reasoning model (2025). Keep only for existing workloads.\n\n```bash\ncurl \"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-pro:generateContent\" \\\n  -H \"x-goog-api-key: $GEMINI_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"contents\":[{\"parts\":[{\"text\":\"Hello\"}]}]}'\n```\n\nSources: [model page](https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro), [pricing](https://ai.google.dev/gemini-api/docs/pricing), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/gemini-2-5-pro/","events_after":881,"major_after":228,"historic_after":45,"missing_at_launch":19,"events_since_release":null,"you_url":"https://postcutoff.com/you/gemini-2-5-pro/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2025-01-xs.md","s":"https://postcutoff.com/briefings/cutoff-2025-01-s.md","m":"https://postcutoff.com/briefings/cutoff-2025-01-m.md","full":"https://postcutoff.com/briefings/cutoff-2025-01.md"}},{"id":"1x-redwood","name":"1X Redwood AI","org":"1X Technologies","family":"Redwood","released":"2025-06-10","status":"current","type":"robotics","modality_in":["image","text","audio"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Onboard 1X NEO (consumer humanoid, preorder)","url":"https://www.1x.tech/order","docs":"https://www.1x.tech/discover/redwood-ai"}],"capabilities":[{"name":"Small onboard VLA for a home humanoid","detail":"160M-parameter vision-language transformer (language embeddings + ViT tokens + proprioception) with a diffusion-policy action decoder, running fully on NEO's embedded GPU at ~5 Hz, so it works without internet.","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/redwood-ai"},{"name":"Mobile bimanual whole-body manipulation","detail":"Combines locomotion with manipulation (bending, leaning, bracing) for retrieving objects, opening doors and navigating the home; learns from both successful and failed episodes.","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/redwood-ai"},{"name":"Voice control via offboard LLM","detail":"An offboard speech-to-speech LLM handles conversation and hands tasks to Redwood.","first":false,"discovered":"launch","source":"https://www.1x.tech/discover/redwood-ai"}],"entry":"2026-04-30-1x-neo-factory","notes":"Ships as NEO's foundational autonomy; tasks it cannot do are handled by remote human teleoperation, which drew privacy criticism (https://startupfortune.com/1xs-20000-neo-robot-lets-a-company-employee-watch-inside-your-home/). NEO: $20,000 Early Access ownership or $499/month subscription, $200 refundable deposit, \"US deliveries start 2026\" (order page checked 2026-09-29). 1X opened its Hayward, CA NEO factory on 2026-04-30 (10,000 units targeted in year one); as of mid-July 2026 no verified customer home delivery had been reported and we found none by 2026-09-29. See also 1x-world-model (video world-model policy, Jan 2026).","verified":"2026-09-29","body_md":"Not callable by developers; only available as the software on a NEO robot.","page_url":"https://postcutoff.com/m/1x-redwood/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":862,"you_url":null,"briefings":null},{"id":"elevenlabs-v3","name":"Eleven v3","org":"ElevenLabs","family":"Eleven v3","released":"2025-06-03","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.08,"unit":"per 1K characters (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.08 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_v3","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/models"},{"provider":"ElevenLabs API (Text to Dialogue)","model_id":"eleven_v3","endpoint":"https://api.elevenlabs.io/v1/text-to-dialogue"},{"provider":"Runway API","model_id":"eleven_v3","endpoint":"https://api.dev.runwayml.com/v1/text_to_speech","docs":"https://docs.dev.runwayml.com/guides/models/"},{"provider":"Web app","url":"https://elevenlabs.io/app"}],"capabilities":[{"name":"Inline audio tags","detail":"Controls delivery with inline tags like [whispers], [laughs], [sighs], [excited]; marketed as 'the most expressive Text to Speech model' at launch. Not marked first: bracketed non-verbal cues existed earlier (e.g. Suno Bark, 2023).","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/eleven-v3"},{"name":"Text to Dialogue (multi-speaker)","detail":"Dedicated Text to Dialogue API for multi-speaker conversations with natural pacing and interruptions.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/eleven-v3"},{"name":"GA release with symbol/number normalization","detail":"GA on 2026-02-02: preferred 72% of the time over alpha; error rate on numbers/symbols/notation cut 68% (15.3% -> 4.9%).","first":false,"discovered":"later","source":"https://elevenlabs.io/blog/eleven-v3-is-now-generally-available"}],"entry":"2026-09-28-elevenlabs-eleven-v4","notes":"Alpha announced 2025-06-03 (blog date); API initially via sales, GA across all platforms 2026-02-02. 70+ languages, 5,000 chars/request. Artificial Analysis TTS arena Elo ~1169 (Sept 2026). Professional Voice Clones not supported on v3 (restored in v4). Real-time variant eleven_v3_conversational has its own file. Voice design: eleven_ttv_v3. Superseded in quality by eleven_v4 (2026-09-28) but still current.","verified":"2026-09-29","body_md":"Expressive TTS with audio tags; superseded in quality by eleven_v4 but still current.\n\n```bash\ncurl -X POST \"https://api.elevenlabs.io/v1/text-to-speech/$VOICE_ID\" -H \"xi-api-key: $ELEVENLABS_API_KEY\" \\\n  -H \"Content-Type: application/json\" -d '{\"text\":\"[whispers] Did you hear that?\",\"model_id\":\"eleven_v3\"}' -o out.mp3\n```\n\n## Changelog\n- 2026-09-29: release date corrected to 2025-06-03 (blog); GA date 2026-02-02, Text to Dialogue, AA Elo added.\n\nSources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/blog/eleven-v3 , https://elevenlabs.io/blog/eleven-v3-is-now-generally-available","page_url":"https://postcutoff.com/m/elevenlabs-v3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":863,"you_url":null,"briefings":null},{"id":"smolvla","name":"SmolVLA (450M)","org":"Hugging Face","family":"LeRobot","released":"2025-06-03","status":"current","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"lerobot/smolvla_base","url":"https://huggingface.co/lerobot/smolvla_base","docs":"https://huggingface.co/docs/lerobot/smolvla"},{"provider":"GitHub (LeRobot)","url":"https://github.com/huggingface/lerobot"}],"capabilities":[{"name":"VLA small enough for a laptop","detail":"450M params (SmolVLM2-500M backbone + flow-matching action expert); trains on a single GPU and runs on consumer hardware incl. MacBooks.","first":false,"discovered":"launch","source":"https://huggingface.co/blog/smolvla"},{"name":"Trained on community-shared data","detail":"Pretrained on ~10M frames from 487 community LeRobot datasets (<30k episodes, an order of magnitude less than other VLAs); 78.3% success on real SO-100 tasks.","first":false,"discovered":"launch","source":"https://huggingface.co/blog/smolvla"},{"name":"Asynchronous inference","detail":"Decouples action prediction from execution: ~30% faster task completion and 2x throughput.","first":false,"discovered":"launch","source":"https://huggingface.co/blog/smolvla"}],"entry":null,"notes":"Designed for low-cost arms (SO-100/SO-101). HF repo still updated in Sept 2026; variants lerobot/smolvla_libero, lerobot/smolvla_robotwin. NVIDIA announced it would acquire Hugging Face (see 2026-09-03 entry).","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/smolvla/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":863,"you_url":null,"briefings":null},{"id":"flux-1-kontext-pro","name":"FLUX.1 Kontext [pro] / [max]","org":"Black Forest Labs","family":"FLUX.1","released":"2025-05-29","status":"legacy","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_image":0.04,"unit":"per image (Kontext pro; Kontext max 0.08)","source":"https://docs.bfl.ai/quick_start/pricing"},"price_line":"$0.04 per image","access":[{"provider":"BFL API","model_id":"flux-kontext-pro","endpoint":"https://api.bfl.ai/v1/flux-kontext-pro","docs":"https://docs.bfl.ai/kontext/kontext_overview"},{"provider":"BFL API","model_id":"flux-kontext-max","endpoint":"https://api.bfl.ai/v1/flux-kontext-max","docs":"https://docs.bfl.ai/kontext/kontext_overview"},{"provider":"Hugging Face (open Kontext [dev])","url":"https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev"}],"capabilities":[{"name":"In-context image editing","detail":"Text-instructed edits of an input image with character consistency across iterative edits; one model for generation and editing.","first":false,"discovered":"launch","source":"https://docs.bfl.ai/kontext/kontext_overview"},{"name":"Open-weight editing sibling","detail":"FLUX.1 Kontext [dev] released as open weights (non-commercial) for local instruction-based editing.","first":false,"discovered":"launch","source":"https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev"}],"entry":null,"notes":"Previous generation (BFL pricing page lists FLUX.1 as 'previous generation'); still served. Release date from memory of BFL launch (May 2025), not re-verified today. Also still served: flux-pro-1.1 ($0.04), flux-pro-1.1-ultra ($0.06).","verified":"2026-09-29","body_md":"Previous-gen FLUX editing models; for new work prefer FLUX.2 [pro]/[max].\n\n```bash\ncurl -X POST https://api.bfl.ai/v1/flux-kontext-pro -H \"x-key: $BFL_API_KEY\" -H \"Content-Type: application/json\" \\\n  -d '{\"prompt\":\"make the car red\",\"input_image\":\"<base64>\"}'\n```\n\nSources: https://docs.bfl.ai/kontext/kontext_overview , https://docs.bfl.ai/quick_start/pricing , https://api.bfl.ai/openapi.json","page_url":"https://postcutoff.com/m/flux-1-kontext-pro/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":863,"you_url":null,"briefings":null},{"id":"hume-evi","name":"Hume EVI 3 / EVI 4 mini (speech-to-speech)","org":"Hume AI","family":"EVI","released":"2025-05-29","status":"current","type":"audio/speech","modality_in":["audio","text"],"modality_out":["audio","text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.06,"unit":"per minute overage on Free/Pro ($0.07 Starter/Creator, $0.05 Scale, $0.04 Business); plans include monthly minutes","source":"https://www.hume.ai/pricing"},"price_line":"$0.06 per minute","access":[{"provider":"Hume API (EVI WebSocket)","model_id":"EVI version 3 or 4-mini (set in EVI config)","endpoint":"wss://api.hume.ai/v0/evi/chat","docs":"https://dev.hume.ai/docs/speech-to-speech-evi/overview"},{"provider":"Web app","url":"https://platform.hume.ai"}],"capabilities":[{"name":"Empathic voice interface with any prompted voice","detail":"EVI 3 (2025-05-29) is a speech-to-speech foundation model that can speak in any of 100,000+ custom voices created via prompting, with inferred personality; ~1.2 s practical end-of-speech-to-response latency at launch.","first":false,"discovered":"launch","source":"https://www.hume.ai/blog/introducing-evi-3"},{"name":"EVI 4 mini: Octave 2 voice in 11 languages","detail":"EVI 4 mini (announced with Octave 2 on 2025-10-01) brings Octave 2 to the speech-to-speech API in 11 languages but must be paired with an external LLM (Anthropic, OpenAI, Google, Fireworks...) until the full EVI 4 ships.","first":false,"discovered":"launch","source":"https://www.hume.ai/blog/octave-2-launch"}],"entry":null,"notes":"EVI 3 is English-only and can answer without an external LLM ('quick responses'); EVI 4 mini is multilingual but requires a supplemental LLM. Both share the same WebSocket; version is chosen in the EVI configuration. Full EVI 4 not launched as of 2026-09-29 (not on Hume blog). EVI 1/2 are older generations.","verified":"2026-09-29","body_md":"Hume's Empathic Voice Interface for real-time voice agents.\n\nSources: https://dev.hume.ai/docs/speech-to-speech-evi/overview , https://www.hume.ai/pricing , https://www.hume.ai/blog/introducing-evi-3","page_url":"https://postcutoff.com/m/hume-evi/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":863,"you_url":null,"briefings":null},{"id":"chatterbox","name":"Resemble AI Chatterbox (Turbo / Nano / Multilingual V3)","org":"Resemble AI","family":"Chatterbox","released":"2025-05-28","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"ResembleAI/chatterbox","url":"https://huggingface.co/ResembleAI/chatterbox"},{"provider":"Hugging Face","model_id":"ResembleAI/chatterbox-turbo","url":"https://huggingface.co/ResembleAI/chatterbox-turbo"},{"provider":"pip","model_id":"chatterbox-tts","url":"https://github.com/resemble-ai/chatterbox"},{"provider":"NVIDIA NIM","model_id":"resembleai/chatterbox-multilingual-tts","url":"https://build.nvidia.com/resembleai/chatterbox-multilingual-tts/modelcard"}],"capabilities":[{"name":"Emotion exaggeration control","detail":"Original 0.5B Chatterbox exposes an exaggeration/intensity knob plus CFG; zero-shot cloning from ~5 s.","first":false,"discovered":"launch","source":"https://github.com/resemble-ai/chatterbox"},{"name":"Chatterbox-Turbo: one-step decoder, paralinguistic tags","detail":"350M params (Dec 2025); speech-token-to-mel decoder distilled from 10 steps to 1; native [laugh], [cough], [chuckle] tags; sub-200 ms production latency.","first":false,"discovered":"later","source":"https://huggingface.co/ResembleAI/chatterbox-turbo"},{"name":"Built-in PerTh watermark","detail":"Every output carries Resemble's imperceptible Perth neural watermark that survives MP3 compression and edits.","first":false,"discovered":"launch","source":"https://github.com/resemble-ai/chatterbox"},{"name":"Multilingual V3 and Nano","detail":"Multilingual V3 (0.5B, 23 languages, better speaker similarity, fewer hallucinations) plus single-language fine-tune packs; Chatterbox-Nano (110M, English, ~3x real time on 8 CPU cores).","first":false,"discovered":"later","source":"https://github.com/resemble-ai/chatterbox"}],"entry":null,"notes":"All MIT-licensed. Multilingual (23 langs) first released Sept 2025. Multilingual V3 released 2026-06-10 (Resemble post; V3 T3 weights first pushed to HF 2026-04-22): same 0.5B Llama backbone, training data up from 25.6k to 36.7k hours, 25 languages incl. 4 dialects and 6 tuned Language Pack models, PerTh watermark on by default; Resemble reports CER under 0.20% for Italian/German but ~71-75% for Korean/Vietnamese (not production-ready); NVIDIA NIM claims 2x-39x throughput. Chatterbox-Nano HF repo created 2026-04-14 (public announcement date not found). Artificial Analysis lists Chatterbox at ~1020 Elo (secondary source). Resemble's pricing page now centres on deepfake detection; hosted TTS price not verified.","verified":"2026-09-29","body_md":"Sources: https://www.resemble.ai/resources/chatterbox-multilingual-v3-tts-with-embedded-watermarking-for-25-languages , https://huggingface.co/ResembleAI/chatterbox-nano , https://github.com/resemble-ai/chatterbox , https://www.resemble.ai/learn/models/chatterbox-multilingual , https://huggingface.co/ResembleAI/chatterbox-turbo","page_url":"https://postcutoff.com/m/chatterbox/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":863,"you_url":null,"briefings":null},{"id":"claude-sonnet-4","name":"Claude Sonnet 4","org":"Anthropic","family":"Claude 4","released":"2025-05-22","status":"retired","type":"reasoning-llm","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":3,"output":15,"unit":"per 1M tokens (USD), Bedrock/Google Cloud may differ","source":"https://platform.claude.com/docs/en/about-claude/pricing"},"price_line":"$3 in, $15 out per 1M tokens","access":[{"provider":"Anthropic API (Claude API)","model_id":"claude-sonnet-4-20250514","endpoint":"https://api.anthropic.com/v1/messages","docs":"https://platform.claude.com/docs/en/about-claude/model-deprecations"},{"provider":"OpenRouter","model_id":"anthropic/claude-sonnet-4","url":"https://openrouter.ai/anthropic/claude-sonnet-4"}],"capabilities":[{"name":"Extended thinking with tool use","detail":"The Claude 4 generation introduced interleaving tool use (e.g. web search) with extended thinking, plus parallel tool calls.","first":true,"discovered":"launch","source":"https://www.anthropic.com/news/claude-4"},{"name":"SOTA SWE-bench at launch","detail":"72.7% on SWE-bench Verified; chosen by GitHub to power the Copilot coding agent.","first":false,"discovered":"launch","source":"https://www.anthropic.com/news/claude-4"}],"entry":null,"notes":"Retired on the Claude API 2026-06-15 (replacement claude-sonnet-4-6 / claude-sonnet-5-5); still available on Amazon Bedrock and Google Cloud per Anthropic pricing page. Cloud ids not re-verified today.","verified":"2026-09-29","body_md":"Retired on Anthropic-operated platforms; requests to `claude-sonnet-4-20250514` on the Claude API fail. Listed for reference because it may still be reachable via Bedrock, Google Cloud or OpenRouter.\n\nSources: https://platform.claude.com/docs/en/about-claude/model-deprecations , https://platform.claude.com/docs/en/about-claude/pricing","page_url":"https://postcutoff.com/m/claude-sonnet-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":864,"you_url":null,"briefings":null},{"id":"gemini-2-5-tts","name":"Gemini 2.5 Flash TTS / Pro TTS","org":"Google DeepMind","family":"Gemini 2.5","released":"2025-05-20","status":"legacy","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.5,"output":10,"unit":"per 1M tokens (USD) for Flash TTS; Pro TTS $1.00 in / $20.00 audio out (25 audio tokens per second)","source":"https://ai.google.dev/gemini-api/docs/pricing"},"price_line":"$0.50 in, $10 out per 1M tokens","access":[{"provider":"Gemini API","model_id":"gemini-2.5-flash-preview-tts","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"Gemini API","model_id":"gemini-2.5-pro-preview-tts","docs":"https://ai.google.dev/gemini-api/docs/models"},{"provider":"Google Cloud Text-to-Speech (Gemini-TTS, GA)","model_id":"gemini-2.5-flash-tts","docs":"https://docs.cloud.google.com/text-to-speech/docs/release-notes"},{"provider":"Google Cloud Text-to-Speech (Gemini-TTS, GA)","model_id":"gemini-2.5-pro-tts"},{"provider":"Google Cloud Text-to-Speech (preview)","model_id":"gemini-2.5-flash-lite-preview-tts"}],"capabilities":[{"name":"Prompt-controlled multi-speaker TTS","detail":"Natural-language control of style, accent, pace and emotion; single and multi-speaker synthesis; 30 speakers in 80+ locales (Cloud GA).","first":false,"discovered":"launch","source":"https://docs.cloud.google.com/text-to-speech/docs/release-notes"}],"entry":null,"notes":"Gemini API ids are 'Limited Access' preview with no shutdown date (migrate to gemini-3.8-flash-tts / -lite-tts). In Cloud TTS, gemini-2.5-flash-tts and gemini-2.5-pro-tts went GA 2025-09-30; streaming added 2025-11-07. Released date = Google I/O 2025 preview (from memory, not re-verified today); Dec 10 2025 update improved expressivity and pacing.","verified":"2026-09-29","body_md":"Sources: [Gemini models](https://ai.google.dev/gemini-api/docs/models), [Gemini pricing](https://ai.google.dev/gemini-api/docs/pricing), [Cloud TTS release notes](https://docs.cloud.google.com/text-to-speech/docs/release-notes), [Cloud TTS pricing](https://cloud.google.com/text-to-speech/pricing).","page_url":"https://postcutoff.com/m/gemini-2-5-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":866,"you_url":null,"briefings":null},{"id":"pi-0-5","name":"π0.5","org":"Physical Intelligence","family":"π (pi)","released":"2025-04-22","status":"current","type":"robotics","modality_in":["text","image"],"modality_out":["action","text"],"open_weights":true,"model_license":"apache-2.0 (openpi code); weights built on PaliGemma, Gemma terms apply","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"GitHub (openpi, JAX + PyTorch)","url":"https://github.com/Physical-Intelligence/openpi","model_id":"gs://openpi-assets/checkpoints/pi05_base"},{"provider":"Hugging Face (LeRobot port)","url":"https://huggingface.co/lerobot/pi05_base","model_id":"lerobot/pi05_base","docs":"https://huggingface.co/docs/lerobot/en/pi05"}],"capabilities":[{"name":"Open-world generalization to unseen homes","detail":"Cleans kitchens and bedrooms in entirely new homes not in training; performance improved as training grew from 3 to 104 homes; co-trained on heterogeneous robot, web and verbal-instruction data.","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pi05"},{"name":"Hierarchical subtask prediction + actions in one model","detail":"Predicts a high-level text subtask, then low-level actions; trained with knowledge insulation.","first":false,"discovered":"launch","source":"https://github.com/Physical-Intelligence/openpi"},{"name":"Newest open-weights π model","detail":"Base plus LIBERO and DROID checkpoints released in openpi in September 2025; the latest PI model with public weights as of 2026-09 (π0.6/π0.7 are closed).","first":false,"discovered":"later","source":"https://github.com/Physical-Intelligence/openpi"}],"entry":null,"notes":"Announced 2025-04-22; weights open-sourced Sept 2025 (pi05_base, pi05_libero, pi05_droid). Still the most capable open-weights PI model. Full fine-tuning needs >70 GB VRAM.","verified":"2026-09-29","body_md":"Sources: [π0.5 blog](https://www.pi.website/blog/pi05), [openpi](https://github.com/Physical-Intelligence/openpi), [LeRobot docs](https://huggingface.co/docs/lerobot/en/pi05).","page_url":"https://postcutoff.com/m/pi-0-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":870,"you_url":null,"briefings":null},{"id":"o3","name":"o3","org":"OpenAI","family":"o-series","released":"2025-04-16","status":"deprecated","type":"reasoning-llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":200000,"max_output":100000,"knowledge_cutoff":"2024-06","pricing":{"input":2,"cached_input":0.5,"output":8,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2 in, $8 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"o3","endpoint":"https://api.openai.com/v1/responses","docs":"https://developers.openai.com/api/docs/models/o3"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"o3","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/o3","url":"https://openrouter.ai/openai/o3"}],"capabilities":[{"name":"Thinking with images","detail":"Reasoning model accepting image input with reasoning tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/o3"},{"name":"Successor: GPT-5","detail":"Docs mark o3 as succeeded by GPT-5; o-series is legacy.","first":false,"discovered":"later","source":"https://developers.openai.com/api/docs/models/o3"}],"entry":null,"notes":"Snapshot o3-2025-04-16 (and o3-pro-2025-06-10) deprecated Jun 11 2026, shutdown Dec 11 2026; replace with gpt-5.6-*. o4-mini-2025-04-16 shuts down Oct 23 2026.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"o3\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/o3\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/o3/","events_after":904,"major_after":242,"historic_after":53,"missing_at_launch":34,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-06.md"}},{"id":"deepgram-aura-2","name":"Deepgram Aura-2","org":"Deepgram","family":"Aura","released":"2025-04-15","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.03,"unit":"per 1K characters PAYG ($0.027 Growth); Aura-1 $0.015","source":"https://deepgram.com/pricing"},"price_line":"$0.03 per 1K characters","access":[{"provider":"Deepgram API","model_id":"aura-2-thalia-en","endpoint":"https://api.deepgram.com/v1/speak?model=aura-2-thalia-en","docs":"https://developers.deepgram.com/docs/tts-models"}],"capabilities":[{"name":"Enterprise TTS with deployable runtime","detail":"Sub-200 ms TTFB, cloud/VPC/on-prem deployment; model id pattern aura-2-{voice}-{lang}.","first":false,"discovered":"launch","source":"https://deepgram.com/learn/introducing-aura-2-enterprise-text-to-speech"},{"name":"7 languages, EN/ES code-switching voices","detail":"English, Spanish, German, French, Dutch, Italian, Japanese; several Spanish voices code-switch with English.","first":false,"discovered":"later","source":"https://developers.deepgram.com/docs/tts-models"}],"entry":null,"notes":"For English voice agents Deepgram now recommends Flux TTS (deepgram-flux-tts); Aura-2 remains the multilingual option.","verified":"2026-09-29","body_md":"Sources: https://developers.deepgram.com/docs/tts-models , https://deepgram.com/pricing","page_url":"https://postcutoff.com/m/deepgram-aura-2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":870,"you_url":null,"briefings":null},{"id":"gpt-4-1","name":"GPT-4.1","org":"OpenAI","family":"GPT-4.1","released":"2025-04-14","status":"legacy","type":"llm","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":1047576,"max_output":32768,"knowledge_cutoff":"2024-06","pricing":{"input":2,"cached_input":0.5,"output":8,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2 in, $8 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-4.1","endpoint":"https://api.openai.com/v1/chat/completions","docs":"https://developers.openai.com/api/docs/models/gpt-4.1"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-4.1","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-4.1","url":"https://openrouter.ai/openai/gpt-4.1"}],"capabilities":[{"name":"1M-token non-reasoning model","detail":"~1M-token context without reasoning tokens; strong instruction following and tool calling.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4.1"},{"name":"Fine-tunable","detail":"Supports fine-tuning, unlike the GPT-5.x models.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4.1"}],"entry":null,"notes":"Snapshot gpt-4.1-2025-04-14. gpt-4.1-mini ($0.40/$1.60) still listed; gpt-4.1-nano deprecated, shutdown Oct 23 2026.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-4.1\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-4.1\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-4-1/","events_after":904,"major_after":242,"historic_after":53,"missing_at_launch":34,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-4-1/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-06.md"}},{"id":"llama-4-maverick","name":"Llama 4 Maverick (17B-128E)","org":"Meta","family":"Llama 4","released":"2025-04-05","status":"legacy","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"llama4-community","context_window":1000000,"max_output":null,"knowledge_cutoff":"2024-08","pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct"},{"provider":"AWS Bedrock","model_id":"meta.llama4-maverick-17b-instruct-v1:0","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-maverick-17b-instruct.html"},{"provider":"OpenRouter","model_id":"meta-llama/llama-4-maverick","url":"https://openrouter.ai/meta-llama/llama-4-maverick"},{"provider":"Web app","url":"https://meta.ai"}],"capabilities":[{"name":"First natively multimodal Llama (early fusion)","detail":"Llama 4 were the first Llama models with native multimodality via early fusion of text and vision tokens.","first":true,"discovered":"launch","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"name":"400B-total MoE on one H100 host","detail":"17B active / 128 experts / ~400B total; runs on a single H100 host.","first":false,"discovered":"launch","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"name":"LMArena experimental-variant controversy","detail":"Launch LMArena Elo 1417 came from an experimental chat-tuned variant, not the released weights, drawing criticism.","first":false,"discovered":"later","source":"https://en.wikipedia.org/wiki/Llama_(language_model)"}],"entry":null,"notes":"FP8 repo meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8. Bedrock max output 8K. No first-party pay-as-you-go pricing verified. Superseded at Meta by closed Muse Spark and open Muse Glimmer.","verified":"2026-09-29","body_md":"Open-weight MoE multimodal model; still widely hosted and cheap on third-party providers.\n\n```bash\ncurl https://openrouter.ai/api/v1/chat/completions -H \"Authorization: Bearer $OPENROUTER_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"meta-llama/llama-4-maverick\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ , https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-maverick-17b-instruct.html","page_url":"https://postcutoff.com/m/llama-4-maverick/","events_after":900,"major_after":239,"historic_after":52,"missing_at_launch":29,"events_since_release":null,"you_url":"https://postcutoff.com/you/llama-4-maverick/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-08-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-08-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-08-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-08.md"}},{"id":"llama-4-scout","name":"Llama 4 Scout (17B-16E)","org":"Meta","family":"Llama 4","released":"2025-04-05","status":"legacy","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":true,"model_license":"llama4-community","context_window":null,"max_output":null,"knowledge_cutoff":"2024-08","pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","url":"https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct"},{"provider":"AWS Bedrock","model_id":"meta.llama4-scout-17b-instruct-v1:0","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-meta-llama-4-scout-17b-instruct.html"},{"provider":"OpenRouter","model_id":"meta-llama/llama-4-scout","url":"https://openrouter.ai/meta-llama/llama-4-scout"},{"provider":"Web app","url":"https://meta.ai"}],"capabilities":[{"name":"10M-token context (claimed)","detail":"Meta advertised an 'industry-leading' 10M-token context via the iRoPE architecture; hosted providers typically serve far less (e.g. ~1.3M on OpenRouter).","first":true,"discovered":"launch","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"},{"name":"Single-H100 multimodal MoE","detail":"17B active / 16 experts / 109B total; fits one H100 with Int4 quantization.","first":false,"discovered":"launch","source":"https://ai.meta.com/blog/llama-4-multimodal-intelligence/"}],"entry":null,"notes":"Context: 10M per Meta; provider limits vary (not listed as context_window). Knowledge cutoff Aug 2024 per Meta model card (not re-verified today). No first-party pricing verified.","verified":"2026-09-29","body_md":"Small open-weight multimodal MoE for long-context and on-prem use.\n\n```bash\ncurl https://openrouter.ai/api/v1/chat/completions -H \"Authorization: Bearer $OPENROUTER_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"meta-llama/llama-4-scout\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ , https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct","page_url":"https://postcutoff.com/m/llama-4-scout/","events_after":900,"major_after":239,"historic_after":52,"missing_at_launch":29,"events_since_release":null,"you_url":"https://postcutoff.com/you/llama-4-scout/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-08-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-08-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-08-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-08.md"}},{"id":"chirp-3-hd","name":"Chirp 3 HD voices (Google Cloud Text-to-Speech)","org":"Google","family":"Chirp 3","released":"2025-04-02","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":30,"unit":"USD per 1M characters (first 1M chars/month free); Chirp 3 Instant custom voice $60 per 1M characters","source":"https://cloud.google.com/text-to-speech/pricing"},"price_line":"$30 per 1M characters","access":[{"provider":"Google Cloud Text-to-Speech API","model_id":"<locale>-Chirp3-HD-<Voice> (e.g. en-US-Chirp3-HD-Charon)","endpoint":"https://texttospeech.googleapis.com (regions global, us, eu, asia-southeast1, europe-west2, asia-northeast1)","docs":"https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd"}],"capabilities":[{"name":"Streaming HD voices in 60+ locales","detail":"28 named voices, streaming and batch synthesis, pace (0.25-2x), pause and IPA/X-SAMPA pronunciation controls, SSML.","first":false,"discovered":"launch","source":"https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd"},{"name":"Instant custom voice","detail":"Chirp 3 Instant custom voice clones a voice from a short sample (30+ locales), priced at $60 per 1M characters.","first":false,"discovered":"later","source":"https://docs.cloud.google.com/text-to-speech/docs/release-notes"}],"entry":null,"notes":"GA 2025-04-02 (8 speakers, 31 locales), since expanded to 60+ locales. Google's enterprise, non-LLM TTS line; the Gemini-TTS models (gemini-2.5-*-tts, Gemini 3.1 Flash TTS) are offered in the same Cloud TTS API. No 2026 successor (e.g. 'Chirp 4') found.","verified":"2026-09-29","body_md":"Sources: [Chirp 3 HD docs](https://docs.cloud.google.com/text-to-speech/docs/chirp3-hd), [Cloud TTS release notes](https://docs.cloud.google.com/text-to-speech/docs/release-notes), [pricing](https://cloud.google.com/text-to-speech/pricing).","page_url":"https://postcutoff.com/m/chirp-3-hd/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":872,"you_url":null,"briefings":null},{"id":"cohere-embed-v4","name":"Cohere Embed v4","org":"Cohere","family":"Embed","released":"2025-04","status":"current","type":"embedding","modality_in":["text","image","pdf"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Cohere API","model_id":"embed-v4.0","endpoint":"https://api.cohere.com/v2/embed","docs":"https://docs.cohere.com/docs/cohere-embed"},{"provider":"AWS Bedrock","model_id":"cohere.embed-v4:0","docs":"https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-cohere-embed-v4.html"}],"capabilities":[{"name":"Interleaved text+image (PDF) embeddings","detail":"Embeds text, images and mixed text/image documents such as PDFs into one vector space.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/cohere-embed"},{"name":"128K-token input with Matryoshka dims","detail":"Up to 128K tokens per input; output dimension selectable 256/512/1024/1536.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/models"}],"entry":null,"notes":"Output is vectors (modality_out text used as placeholder). Release month (Apr 2025) from memory, not re-verified. Pricing not verified (Cohere pricing page shows only Model Vault hourly rates).","verified":"2026-09-29","body_md":"Multimodal embeddings for search/RAG over long docs and scanned PDFs.\n\n```bash\ncurl https://api.cohere.com/v2/embed -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"embed-v4.0\",\"input_type\":\"search_document\",\"embedding_types\":[\"float\"],\"texts\":[\"hello world\"]}'\n```\n\nSources: https://docs.cohere.com/docs/models , https://docs.cohere.com/docs/cohere-embed","page_url":"https://postcutoff.com/m/cohere-embed-v4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":870,"you_url":null,"briefings":null},{"id":"gpt-4o-mini-tts","name":"GPT-4o mini TTS","org":"OpenAI","family":"GPT-4o","released":"2025-03-20","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"text_input":0.6,"audio_output":12,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/models/gpt-4o-mini-tts"},"price_line":"$0.60 text in, $12 audio out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-4o-mini-tts","endpoint":"https://api.openai.com/v1/audio/speech","docs":"https://developers.openai.com/api/docs/models/gpt-4o-mini-tts"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-4o-mini-tts","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Steerable speech","detail":"Only current OpenAI TTS model listed in the models catalog; max 2000 input tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4o-mini-tts"},{"name":"Instruction-steerable voice","detail":"An `instructions` field controls accent, emotional range, intonation, impressions, speed, tone and whispering.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/guides/text-to-speech"}],"entry":null,"notes":"Snapshots gpt-4o-mini-tts-2025-03-20 and gpt-4o-mini-tts-2025-12-15 (default). Older tts-1 ($15/1M chars) and tts-1-hd ($30) still priced.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/audio/speech \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-4o-mini-tts\",\"voice\":\"alloy\",\"input\":\"Hello\"}' --output out.mp3\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-4o-mini-tts\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-4o-mini-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":873,"you_url":null,"briefings":null},{"id":"gpt-4o-transcribe","name":"GPT-4o Transcribe / Mini Transcribe / Transcribe Diarize","org":"OpenAI","family":"GPT-4o","released":"2025-03-20","status":"deprecated","type":"audio/speech","modality_in":["audio","text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":16000,"max_output":2000,"knowledge_cutoff":"2024-06","pricing":{"input":2.5,"output":10,"unit":"per 1M tokens (USD) for gpt-4o-transcribe and gpt-4o-transcribe-diarize (~$0.006/min); gpt-4o-mini-transcribe $1.25 / $5 (~$0.003/min)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2.50 in, $10 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-4o-transcribe","endpoint":"https://api.openai.com/v1/audio/transcriptions","docs":"https://developers.openai.com/api/docs/models/gpt-4o-transcribe"},{"provider":"OpenAI API","model_id":"gpt-4o-mini-transcribe","endpoint":"https://api.openai.com/v1/audio/transcriptions","docs":"https://developers.openai.com/api/docs/models/gpt-4o-mini-transcribe"},{"provider":"OpenAI API","model_id":"gpt-4o-transcribe-diarize","endpoint":"https://api.openai.com/v1/audio/transcriptions"}],"capabilities":[{"name":"LLM-based transcription","detail":"Uses GPT-4o for speech-to-text with better accuracy than the original Whisper models; also usable in Realtime transcription sessions.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4o-transcribe"}],"entry":null,"notes":"Deprecated 2026-08-26, shutdown 2027-02-26 (with whisper-1); replacements gpt-transcribe (files) and gpt-live-transcribe (streaming). gpt-4o-mini-transcribe-2025-03-20 was separately deprecated 2026-07-20 in favour of the 2025-12-15 snapshot. Release date 2025-03-20 is the date of the gpt-4o-mini-tts/transcribe snapshots, not re-verified on an OpenAI launch post.","verified":"2026-09-29","body_md":"Migrate to [gpt-transcribe](gpt-transcribe.md) or [gpt-live-transcribe](gpt-live-transcribe.md).\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-4o-transcribe\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/deprecations","page_url":"https://postcutoff.com/m/gpt-4o-transcribe/","events_after":904,"major_after":242,"historic_after":53,"missing_at_launch":31,"events_since_release":null,"you_url":null,"briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2024-06-xs.md","s":"https://postcutoff.com/briefings/cutoff-2024-06-s.md","m":"https://postcutoff.com/briefings/cutoff-2024-06-m.md","full":"https://postcutoff.com/briefings/cutoff-2024-06.md"}},{"id":"gr00t-n1","name":"Isaac GR00T N1 / N1.5 / N1.6","org":"NVIDIA","family":"Isaac GR00T","released":"2025-03-18","status":"legacy","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":true,"model_license":"NVIDIA license (see each Hugging Face model card; N1.6 card lists a non-commercial license)","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"nvidia/GR00T-N1-2B","url":"https://huggingface.co/nvidia/GR00T-N1-2B"},{"provider":"Hugging Face","model_id":"nvidia/GR00T-N1.5-3B","url":"https://huggingface.co/nvidia/GR00T-N1.5-3B"},{"provider":"Hugging Face","model_id":"nvidia/GR00T-N1.6-3B","url":"https://huggingface.co/nvidia/GR00T-N1.6-3B"},{"provider":"GitHub (branches n1d5, n1d6)","url":"https://github.com/NVIDIA/Isaac-GR00T"}],"capabilities":[{"name":"Open humanoid robot foundation model","detail":"Announced at GTC 2025 as 'the world's first open humanoid robot foundation model': a dual-system VLA (VLM 'System 2' + diffusion-transformer 'System 1') for cross-embodiment humanoid control, customizable with synthetic data.","first":true,"discovered":"launch","source":"https://nvidianews.nvidia.com/news/nvidia-isaac-gr00t-n1-open-humanoid-robot-foundation-model-simulation-frameworks"}],"entry":null,"notes":"N1 (2B) announced 2025-03-18; N1.5 (3B) mid-2025; N1.6 (3B) later in 2025 - exact N1.5/N1.6 dates not re-verified. Superseded by GR00T N1.7 (2026). 'first' is NVIDIA's claim (open weights for a humanoid-specific generalist model; earlier open VLAs such as OpenVLA/Octo targeted arms).","verified":"2026-09-29","body_md":"Sources: [NVIDIA newsroom (GR00T N1)](https://nvidianews.nvidia.com/news/nvidia-isaac-gr00t-n1-open-humanoid-robot-foundation-model-simulation-frameworks), [Isaac-GR00T GitHub](https://github.com/NVIDIA/Isaac-GR00T).","page_url":"https://postcutoff.com/m/gr00t-n1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":873,"you_url":null,"briefings":null},{"id":"sesame-csm-1b","name":"Sesame CSM-1B (Conversational Speech Model)","org":"Sesame","family":"CSM","released":"2025-03-13","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"sesame/csm-1b","url":"https://huggingface.co/sesame/csm-1b"},{"provider":"Transformers","model_id":"sesame/csm-1b","docs":"https://huggingface.co/docs/transformers/model_doc/csm"},{"provider":"Sesame app (Maya, Miles, Simone, Charlie — hosted larger models)","url":"https://www.sesame.com/"}],"capabilities":[{"name":"Context-conditioned conversational TTS","detail":"Llama backbone + audio decoder emitting Mimi audio codes; generates speech conditioned on prior conversation audio/text so prosody fits the dialogue; voice prompting via context segments.","first":false,"discovered":"launch","source":"https://huggingface.co/sesame/csm-1b"}],"entry":"2026-05-28-sesame-ios-app","notes":"Open base generation model only (no fine-tuned voices, English-centric, cannot generate text itself); the Maya/Miles demo voices use Sesame's larger in-house models. Native in Transformers since v4.52.1. Sesame raised a $250M Series B (Oct 2025, Sequoia/Spark) and launched a public-preview iOS app with four agents (Maya, Miles, Simone, Charlie) in 39 countries on 2026-05-28; smart glasses targeted for 2027. No newer open Sesame model found as of 2026-09-29.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/sesame/csm-1b , https://techcrunch.com/2026/05/28/sesame-the-conversational-ai-startup-from-oculus-founders-launches-its-ios-app/","page_url":"https://postcutoff.com/m/sesame-csm-1b/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":873,"you_url":null,"briefings":null},{"id":"agibot-go-1","name":"AgiBot GO-1 (Genie Operator-1)","org":"AgiBot","family":"Genie Operator","released":"2025-03-10","status":"legacy","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"cc-by-nc-sa-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"agibot-world/GO-1","url":"https://huggingface.co/agibot-world/GO-1"},{"provider":"Hugging Face (lighter variant)","model_id":"agibot-world/GO-1-Air","url":"https://huggingface.co/agibot-world/GO-1-Air"},{"provider":"GitHub","url":"https://github.com/OpenDriveLab/Agibot-World"}],"capabilities":[{"name":"Latent-action VLA trained on AgiBot World","detail":"3B model on an InternVL2.5-2B backbone using latent action representations, pretrained on AgiBot World (1M+ trajectories, 217 tasks, 5 deployment scenarios); ~30% average gain over policies trained on Open X-Embodiment, 60%+ success on complex tasks, +32% vs RDT.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2503.06669"}],"entry":null,"notes":"Paper arXiv 2503.06669 (2025-03-09); announced ~2025-03-10 (day not re-verified). Weights on HF from Sept 2025, non-commercial license. Successor: agibot-go-2 (Apr 2026).","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/agibot-go-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":874,"you_url":null,"briefings":null},{"id":"orpheus-tts","name":"Canopy Labs Orpheus TTS (3B)","org":"Canopy Labs","family":"Orpheus","released":"2025-03","status":"current","type":"audio/speech","modality_in":["text","audio"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"canopylabs/orpheus-tts-0.1-finetune-prod","url":"https://github.com/canopyai/Orpheus-TTS"},{"provider":"Hugging Face","model_id":"canopylabs/orpheus-3b-0.1-ft","url":"https://huggingface.co/canopylabs/orpheus-3b-0.1-ft"},{"provider":"Groq","model_id":"canopylabs/orpheus-v1-english","endpoint":"https://api.groq.com/openai/v1/audio/speech","docs":"https://console.groq.com/docs/text-to-speech"},{"provider":"Groq","model_id":"canopylabs/orpheus-arabic-saudi","endpoint":"https://api.groq.com/openai/v1/audio/speech"},{"provider":"Together AI","url":"https://www.together.ai/models/orpheus-tts"}],"capabilities":[{"name":"LLM-backbone TTS with emotion tags","detail":"Llama-3B-based speech LLM trained on 100k+ h English; tags <laugh>, <chuckle>, <sigh>, <cough>, <sniffle>, <groan>, <yawn>, <gasp>; ~200 ms streaming latency (~100 ms with input streaming); zero-shot cloning via pretrained model.","first":false,"discovered":"launch","source":"https://github.com/canopyai/Orpheus-TTS"}],"entry":null,"notes":"8 English preset voices (tara, leah, jess, leo, dan, mia, zac, zoe); multilingual research release (7 language pairs) April 2025. Groq deployed two variants on 2026-01-13 (press: $22 per 1M characters, not verified on Groq pricing page).","verified":"2026-09-29","body_md":"Sources: https://github.com/canopyai/Orpheus-TTS , https://console.groq.com/docs/text-to-speech","page_url":"https://postcutoff.com/m/orpheus-tts/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":872,"you_url":null,"briefings":null},{"id":"command-a","name":"Command A (03-2025) and variants","org":"Cohere","family":"Command A","released":"2025-03","status":"legacy","type":"llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"cc-by-nc-4.0","context_window":256000,"max_output":8000,"knowledge_cutoff":null,"pricing":{"input":2.5,"output":10,"unit":"per 1M tokens (USD) on OpenRouter; Cohere first-party price not verified","source":"https://openrouter.ai/cohere/command-a"},"price_line":"$2.50 in, $10 out per 1M tokens","access":[{"provider":"Cohere API","model_id":"command-a-03-2025","endpoint":"https://api.cohere.com/v2/chat","docs":"https://docs.cohere.com/docs/command-a"},{"provider":"OpenRouter","model_id":"cohere/command-a","url":"https://openrouter.ai/cohere/command-a"},{"provider":"Hugging Face","url":"https://huggingface.co/CohereLabs/c4ai-command-a-03-2025"}],"capabilities":[{"name":"Enterprise model on two GPUs","detail":"111B model that runs on only two A100/H100 GPUs, 150% higher throughput than Command R+ 08-2024.","first":false,"discovered":"launch","source":"https://docs.cohere.com/docs/command-a"},{"name":"Specialized variants","detail":"Separate ids command-a-reasoning-08-2025 (256K/32K out), command-a-vision-07-2025 (image input) and command-a-translate-08-2025 (23-language MT).","first":false,"discovered":"later","source":"https://docs.cohere.com/docs/models"}],"entry":null,"notes":"Superseded by Command A+ (May 2026). Variants listed in notes/capabilities share this file. Weights are non-commercial (CC-BY-NC).","verified":"2026-09-29","body_md":"Tool use, RAG and multilingual agents; still available but Command A+ is newer.\n\n```bash\ncurl https://api.cohere.com/v2/chat -H \"Authorization: Bearer $CO_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"command-a-03-2025\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}'\n```\n\nSources: https://docs.cohere.com/docs/models , https://docs.cohere.com/docs/command-a","page_url":"https://postcutoff.com/m/command-a/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":872,"you_url":null,"briefings":null},{"id":"elevenlabs-scribe-v1","name":"Scribe v1","org":"ElevenLabs","family":"Scribe","released":"2025-02-26","status":"deprecated","type":"audio/speech","modality_in":["audio","video"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"ElevenLabs API","model_id":"scribe_v1","endpoint":"https://api.elevenlabs.io/v1/speech-to-text","docs":"https://elevenlabs.io/docs/models"}],"capabilities":[{"name":"ElevenLabs' first speech-to-text model","detail":"99 languages, word timestamps, diarization and audio-event tagging; claimed highest benchmark accuracy vs Gemini 2.0 and Whisper v3 at launch ($0.40/hour).","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/meet-scribe"}],"entry":null,"notes":"Deprecated on the models page ('First generation speech recognition (outclassed by v2)'). Use scribe_v2. Current price not listed separately.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/blog/meet-scribe","page_url":"https://postcutoff.com/m/elevenlabs-scribe-v1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":874,"you_url":null,"briefings":null},{"id":"figure-helix","name":"Helix (Figure, v1)","org":"Figure AI","family":"Helix","released":"2025-02-20","status":"legacy","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"None (runs only on Figure robots; no public API or weights)","url":"https://www.figure.ai/news/helix"}],"capabilities":[{"name":"Full upper-body humanoid control from a VLA","detail":"Continuous high-rate control of the whole humanoid upper body (wrists, torso, head, individual fingers) over a 35-DoF action space.","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix"},{"name":"Dual-system architecture (S2 + S1)","detail":"System 2: 7B VLM at 7-9 Hz for scene/language understanding; System 1: 80M-parameter visuomotor transformer at 200 Hz.","first":false,"discovered":"launch","source":"https://www.figure.ai/news/helix"},{"name":"Multi-robot collaboration with one set of weights","detail":"Same model ran simultaneously on two robots collaborating on a shared grocery-storage task.","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix"},{"name":"Fully onboard on embedded low-power GPUs","detail":"Runs entirely on the robot's embedded GPUs - Figure calls it the first VLA ready for commercial deployment this way.","first":true,"discovered":"launch","source":"https://www.figure.ai/news/helix"}],"entry":null,"notes":"'first' flags are Figure's own claims at announcement (2025-02-20). Superseded by Helix 02 (2026-01) and Helix 2.5 (2026-09). Never publicly available.","verified":null,"body_md":"Sources: [Figure: Helix](https://www.figure.ai/news/helix).","page_url":"https://postcutoff.com/m/figure-helix/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":876,"you_url":null,"briefings":null},{"id":"deepgram-nova-3","name":"Deepgram Nova-3 (incl. Medical / Pharma)","org":"Deepgram","family":"Nova","released":"2025-02-12","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.0048,"unit":"per minute streaming monolingual PAYG; multilingual $0.0058; pre-recorded $0.0043 mono / $0.0052 multi","source":"https://deepgram.com/pricing"},"price_line":"$0.0048 per minute","access":[{"provider":"Deepgram API","model_id":"nova-3","endpoint":"https://api.deepgram.com/v1/listen (REST) / wss://api.deepgram.com/v1/listen","docs":"https://developers.deepgram.com/docs/models-languages-overview"},{"provider":"Deepgram API","model_id":"nova-3-medical"},{"provider":"Deepgram API","model_id":"nova-3-pharma"}],"capabilities":[{"name":"Keyterm prompting, 90+ languages","detail":"Nova-3 general supports 90+ languages incl. multilingual code-switching mode; languages added continuously through 2026 (e.g. Kazakh 2026-09-03, Assamese/Mongolian/Pashto 2026-08-27).","first":false,"discovered":"later","source":"https://developers.deepgram.com/changelog"},{"name":"Domain variants","detail":"nova-3-medical (upgraded batch model May 2026) and nova-3-pharma (English pharmaceutical model, 2026-09-17).","first":false,"discovered":"later","source":"https://developers.deepgram.com/changelog"}],"entry":null,"notes":"Release date 2025-02-12 from Deepgram's Nova-3 launch (not re-checked today). Previous gen nova-2 and variants still served. Deepgram also hosts Whisper (whisper-large, $0.0048/min).","verified":"2026-09-29","body_md":"Deepgram's general-purpose transcription model (batch + streaming).\n\nSources: https://developers.deepgram.com/docs/models-languages-overview , https://deepgram.com/pricing , https://developers.deepgram.com/changelog","page_url":"https://postcutoff.com/m/deepgram-nova-3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":878,"you_url":null,"briefings":null},{"id":"kokoro-82m","name":"Kokoro-82M","org":"hexgrad","family":"Kokoro","released":"2025-01-27","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":true,"model_license":"apache-2.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":0.65,"unit":"USD per 1M characters (cheapest hosted price per Artificial Analysis); self-hosting free","source":"https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice"},"price_line":"$0.65 per 1M characters","access":[{"provider":"Hugging Face","model_id":"hexgrad/Kokoro-82M","url":"https://huggingface.co/hexgrad/Kokoro-82M"},{"provider":"pip","model_id":"kokoro","url":"https://github.com/hexgrad/kokoro"},{"provider":"DeepInfra","model_id":"hexgrad/Kokoro-82M","url":"https://deepinfra.com/hexgrad/Kokoro-82M"},{"provider":"OpenRouter","url":"https://openrouter.ai/hexgrad/kokoro-82m"}],"capabilities":[{"name":"Tiny model, top-tier quality","detail":"82M-param StyleTTS2 + ISTFTNet model trained for ~$1,000 (1,000 A100 h) on permissive data; v0.19 hit #1 on the HF TTS Spaces Arena; still top-5 open weights on Artificial Analysis (~1065 Elo) in Sept 2026.","first":false,"discovered":"launch","source":"https://huggingface.co/hexgrad/Kokoro-82M"}],"entry":null,"notes":"v1.0: 54 preset voices, 8 languages (US/UK English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, Mandarin), 24 kHz. No voice cloning. v0.19 was 2024-12-25. ~11.5M HF downloads/month.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/hexgrad/Kokoro-82M , https://artificialanalysis.ai/text-to-speech/leaderboard/provider-voice","page_url":"https://postcutoff.com/m/kokoro-82m/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":881,"you_url":null,"briefings":null},{"id":"pi-0-fast","name":"π0-FAST","org":"Physical Intelligence","family":"π (pi)","released":"2025-01-16","status":"legacy","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0 (openpi code); weights built on PaliGemma, Gemma terms apply","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"GitHub (openpi)","url":"https://github.com/Physical-Intelligence/openpi","model_id":"gs://openpi-assets/checkpoints/pi0_fast_base"},{"provider":"Hugging Face (LeRobot port)","url":"https://huggingface.co/lerobot/pi0fast-base","model_id":"lerobot/pi0fast-base","docs":"https://huggingface.co/docs/lerobot/pi0fast"}],"capabilities":[{"name":"FAST action tokenizer (autoregressive VLA)","detail":"Frequency-space Action Sequence Tokenization (DCT + BPE) compresses action chunks ~10x, letting an autoregressive VLA learn dexterous high-frequency tasks and train up to 5x faster than diffusion/flow π0.","first":false,"discovered":"launch","source":"https://huggingface.co/blog/pi0"},{"name":"DROID generalist checkpoint","detail":"pi0_fast_droid runs zero-shot on Franka DROID setups for many table-top instructions (openpi).","first":false,"discovered":"later","source":"https://github.com/Physical-Intelligence/openpi"}],"entry":null,"notes":"FAST tokenizer released and open-sourced mid-January 2025 (X post by @physical_int, 2025-01-16 approx.); π0-FAST weights open-sourced in openpi on 2025-02-04. Not marked first: no explicit 'first' claim verified.","verified":"2026-09-29","body_md":"Autoregressive sibling of π0 built on the FAST tokenizer.\n\nSources: [HF blog: π0 and π0-FAST](https://huggingface.co/blog/pi0), [PI on X](https://x.com/physical_int/status/1879963467836453067), [openpi](https://github.com/Physical-Intelligence/openpi).","page_url":"https://postcutoff.com/m/pi-0-fast/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":884,"you_url":null,"briefings":null},{"id":"elevenlabs-flash-v2-5","name":"Eleven Flash v2.5 / Flash v2","org":"ElevenLabs","family":"Eleven Flash","released":"2024-12-18","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.04,"unit":"per 1K characters (API, Flash/Turbo tier)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.04 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_flash_v2_5","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/models"},{"provider":"ElevenLabs API (English only)","model_id":"eleven_flash_v2","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/models"}],"capabilities":[{"name":"~75 ms TTS","detail":"Ultra-fast model for real-time use: ~75 ms model latency (excl. application & network). Flash v2.5: 32 languages, 40,000 chars/request; Flash v2: English only, 30,000 chars.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"}],"entry":null,"notes":"Announced 2024-12-18 ('Meet Flash', X post). Replaced Turbo v2/v2.5 (now deprecated). Text normalization available for Flash v2.5 (enterprise). For expressive real-time use ElevenLabs now points to eleven_v4_turbo (~100 ms).","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api , https://elevenlabs.io/blog/meet-flash , https://x.com/ElevenLabs/status/1869462840941461941","page_url":"https://postcutoff.com/m/elevenlabs-flash-v2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":889,"you_url":null,"briefings":null},{"id":"phi-4","name":"Phi-4 (14B)","org":"Microsoft","family":"Phi-4","released":"2024-12-12","status":"legacy","type":"llm","modality_in":["text"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":16384,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.07,"output":0.14,"unit":"per 1M tokens (USD) on OpenRouter; self-hosting free","source":"https://openrouter.ai/microsoft/phi-4"},"price_line":"$0.07 in, $0.14 out per 1M tokens","access":[{"provider":"Hugging Face","url":"https://huggingface.co/microsoft/phi-4"},{"provider":"Azure AI Foundry","model_id":"Phi-4","url":"https://azure.microsoft.com/en-us/products/phi"},{"provider":"OpenRouter","model_id":"microsoft/phi-4","url":"https://openrouter.ai/microsoft/phi-4"}],"capabilities":[{"name":"Synthetic-data small model","detail":"14B dense model trained on 9.8T tokens heavy in curated synthetic data, prioritizing reasoning over scale (84.8 MMLU, 80.4 MATH).","first":false,"discovered":"launch","source":"https://huggingface.co/microsoft/phi-4"},{"name":"Reasoning derivatives","detail":"Base for Phi-4-reasoning, Phi-4-reasoning-plus, Phi-4-mini(-reasoning/-flash-reasoning) and Phi-4-multimodal-instruct open models.","first":false,"discovered":"later","source":"https://huggingface.co/microsoft"}],"entry":null,"notes":"No Phi-5 found on Hugging Face as of 2026-09-29 (microsoft org). Foundry model name not re-verified today. Siblings: microsoft/Phi-4-mini-instruct, microsoft/Phi-4-reasoning-plus, microsoft/Phi-4-multimodal-instruct.","verified":"2026-09-29","body_md":"Small open model for local/edge reasoning, math and English text tasks.\n\n```python\nfrom transformers import pipeline\npipe = pipeline(\"text-generation\", model=\"microsoft/phi-4\", device_map=\"auto\")\nprint(pipe([{\"role\":\"user\",\"content\":\"Solve 2x+3=11\"}], max_new_tokens=128))\n```\n\nSources: https://huggingface.co/microsoft/phi-4 , https://azure.microsoft.com/en-us/products/phi","page_url":"https://postcutoff.com/m/phi-4/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":889,"you_url":null,"briefings":null},{"id":"pi-0","name":"π0 (pi-zero)","org":"Physical Intelligence","family":"π (pi)","released":"2024-10-31","status":"legacy","type":"robotics","modality_in":["text","image"],"modality_out":["action"],"open_weights":true,"model_license":"apache-2.0 (openpi code); weights built on PaliGemma, Gemma terms apply","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"GitHub (openpi, JAX + PyTorch)","url":"https://github.com/Physical-Intelligence/openpi","model_id":"gs://openpi-assets/checkpoints/pi0_base"},{"provider":"Hugging Face (LeRobot port)","url":"https://huggingface.co/lerobot/pi0_base","model_id":"lerobot/pi0_base","docs":"https://huggingface.co/docs/lerobot/pi0"}],"capabilities":[{"name":"Flow-matching VLA for dexterous, high-frequency control","detail":"PaliGemma VLM plus an action expert that outputs continuous action chunks via flow matching (up to 50 Hz), trained on data from 8 distinct robots; folds laundry, busses tables, assembles boxes.","first":false,"discovered":"launch","source":"https://www.pi.website/blog/pi0"},{"name":"Open weights with fine-tuning recipes","detail":"Open-sourced 2025-02-04 in openpi with base and fine-tuned checkpoints (ALOHA towel/tupperware/pen, DROID) pre-trained on 10k+ hours of robot data; inference needs >8 GB VRAM, LoRA fine-tuning >22.5 GB.","first":false,"discovered":"later","source":"https://github.com/Physical-Intelligence/openpi"}],"entry":null,"notes":"Announced 2024-10-31; weights released 2025-02-04 (openpi). Not the first open VLA (OpenVLA/Octo came earlier) but became the most widely used open generalist robot policy baseline. Superseded by π0.5; still available. Fine-tuned expert checkpoints: pi0_droid, pi0_aloha_towel, pi0_aloha_tupperware, pi0_aloha_pen_uncap.","verified":"2026-09-29","body_md":"Physical Intelligence's first generalist robot policy and the base of the open-source openpi stack.\n\nSources: [π0 blog](https://www.pi.website/blog/pi0), [Open Sourcing π0](https://www.pi.website/blog/openpi), [openpi](https://github.com/Physical-Intelligence/openpi), [The Robot Report](https://www.therobotreport.com/physical-intelligence-open-sources-pi0-robotics-foundation-model/).","page_url":"https://postcutoff.com/m/pi-0/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":893,"you_url":null,"briefings":null},{"id":"stable-diffusion-3-5-large","name":"Stable Diffusion 3.5 Large","org":"Stability AI","family":"Stable Diffusion 3.5","released":"2024-10-22","status":"current","type":"image-gen","modality_in":["text","image"],"modality_out":["image"],"open_weights":true,"model_license":"Stability AI Community License","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Stability AI API","model_id":"sd3.5-large","endpoint":"https://api.stability.ai/v2beta/stable-image/generate/sd3","docs":"https://platform.stability.ai/docs/api-reference"},{"provider":"Hugging Face","url":"https://huggingface.co/stabilityai/stable-diffusion-3.5-large"},{"provider":"Hugging Face (Large Turbo)","url":"https://huggingface.co/stabilityai/stable-diffusion-3.5-large-turbo"},{"provider":"Hugging Face (Medium)","url":"https://huggingface.co/stabilityai/stable-diffusion-3.5-medium"}],"capabilities":[{"name":"Open MMDiT weights, free for small businesses","detail":"8B Multimodal Diffusion Transformer with open weights under a Community License free for commercial use under $1M annual revenue.","first":false,"discovered":"launch","source":"https://huggingface.co/stabilityai/stable-diffusion-3.5-large"},{"name":"Broad hardware optimization","detail":"Official TensorRT/FP8 (NVIDIA, ~2x faster, 40% less memory), ONNX AMD GPU and AMD NPU builds released later.","first":false,"discovered":"later","source":"https://stability.ai/news-updates"}],"entry":null,"notes":"Still Stability's latest image model family (no official SD4 as of 2026-09; SD4 'news' articles are unverified). API model values for /generate/sd3: sd3.5-large, sd3.5-large-turbo, sd3.5-medium (from third-party docs; official API ref is JS-rendered, not verified). Pricing (credits) not verified.","verified":null,"body_md":"Main open-weight Stability image model; widely used with ComfyUI/diffusers and fine-tunes.\n\n```bash\ncurl -X POST https://api.stability.ai/v2beta/stable-image/generate/sd3 -H \"Authorization: Bearer $STABILITY_API_KEY\" -H \"Accept: image/*\" \\\n  -F prompt=\"a lighthouse at dusk, oil painting\" -F model=sd3.5-large -o out.png\n```\n\nSources: https://huggingface.co/stabilityai/stable-diffusion-3.5-large , https://stability.ai/news-updates , https://platform.stability.ai/docs/api-reference","page_url":"https://postcutoff.com/m/stable-diffusion-3-5-large/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":893,"you_url":null,"briefings":null},{"id":"kyutai-moshi","name":"Kyutai Moshi / Hibiki-Zero (full-duplex speech models)","org":"Kyutai","family":"Moshi","released":"2024-09-17","status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["audio","text"],"open_weights":true,"model_license":"cc-by-4.0","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"kyutai/moshiko-pytorch-bf16","url":"https://github.com/kyutai-labs/moshi"},{"provider":"Hugging Face","model_id":"kyutai/hibiki-zero-3b-pytorch-bf16","url":"https://huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16"},{"provider":"Web demo","url":"https://moshi.chat"}],"capabilities":[{"name":"Open full-duplex spoken dialogue","detail":"7B temporal transformer modelling user and Moshi audio streams simultaneously with an 'inner monologue' text stream; 160 ms theoretical / ~200 ms practical latency on an L4; Mimi codec (24 kHz, 12.5 Hz, 1.1 kbps).","first":true,"discovered":"launch","source":"https://github.com/kyutai-labs/moshi"},{"name":"Hibiki-Zero simultaneous speech translation","detail":"3B model (2026-02-12) translating French, Spanish, Portuguese and German speech to English in real time with voice transfer, trained without aligned data.","first":false,"discovered":"later","source":"https://kyutai.org/blog/"},{"name":"MoshiRAG","detail":"Asynchronous knowledge retrieval via a text LLM for full-duplex speech models (2026-04-30); RL post-training for interactivity (2026-06-10).","first":false,"discovered":"later","source":"https://kyutai.org/blog/"}],"entry":null,"notes":"Moshi (announced July 2024, weights + paper Sept 2024) is widely cited as the first real-time full-duplex open spoken dialogue model; NVIDIA PersonaPlex-7B (Jan 2026) is fine-tuned from Moshiko weights. Variants: moshiko (male)/moshika (female) in PyTorch bf16/int8, MLX int4/int8/bf16, Rust/Candle. Code MIT/Apache, weights CC-BY-4.0.","verified":"2026-09-29","body_md":"Sources: https://github.com/kyutai-labs/moshi , https://kyutai.org/blog/ , https://huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16","page_url":"https://postcutoff.com/m/kyutai-moshi/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":898,"you_url":null,"briefings":null},{"id":"openvla","name":"OpenVLA (7B) and OpenVLA-OFT","org":"Stanford / UC Berkeley / Toyota Research Institute","family":"OpenVLA","released":"2024-06-13","status":"legacy","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"openvla/openvla-7b","url":"https://huggingface.co/openvla/openvla-7b"},{"provider":"Hugging Face (OFT fine-tunes)","model_id":"moojink/openvla-7b-oft-finetuned-libero-spatial","url":"https://huggingface.co/moojink/openvla-7b-oft-finetuned-libero-spatial"},{"provider":"GitHub","url":"https://github.com/openvla/openvla","docs":"https://openvla.github.io/"}],"capabilities":[{"name":"Open 7B generalist VLA beating a 55B closed model","detail":"Llama 2 7B backbone with fused DINOv2 + SigLIP vision, trained on ~970k Open X-Embodiment episodes; outperformed RT-2-X (55B) by 16.5% absolute success over 29 tasks with 7x fewer parameters, and fine-tunes with LoRA on consumer GPUs.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2406.09246"},{"name":"OFT fine-tuning recipe (Feb 2025)","detail":"OpenVLA-OFT (parallel decoding, action chunking, continuous actions, L1 loss) raised LIBERO average success from 76.5% to 97.1% and action throughput 26x; on bimanual ALOHA it beat pi0 and RDT-1B by up to 15% absolute.","first":false,"discovered":"later","source":"https://arxiv.org/abs/2502.19645"}],"entry":null,"notes":"The most-downloaded open VLA checkpoint on HF (500k+ downloads at check time); widely used as a research baseline. Superseded in capability by pi0-family and newer open VLAs but still a standard reference. Release day: arXiv 2406.09246 v1 dated 2024-06-13 (HF repo created 2024-06-10).","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/openvla/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":906,"you_url":null,"briefings":null},{"id":"octo","name":"Octo (Octo-Small / Octo-Base 1.5)","org":"UC Berkeley (RAIL) / Stanford / CMU / Google DeepMind","family":"Octo","released":"2024-05-20","status":"legacy","type":"robotics","modality_in":["image","text"],"modality_out":["action"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"rail-berkeley/octo-base-1.5","url":"https://huggingface.co/rail-berkeley/octo-base-1.5"},{"provider":"GitHub","url":"https://github.com/octo-models/octo","docs":"https://octo-models.github.io/"}],"capabilities":[{"name":"Open generalist policy on Open X-Embodiment","detail":"Transformer diffusion policy (27M Small / 93M Base) trained on 800k trajectories from Open X-Embodiment; instructed by language or goal images; evaluated on 9 robot platforms; fine-tunes to new sensors and action spaces in hours on consumer GPUs.","first":false,"discovered":"launch","source":"https://arxiv.org/abs/2405.12213"}],"entry":null,"notes":"Early (2024) fully open generalist robot policy; now mostly a baseline. Parameter sizes from the project page.","verified":"2026-09-29","body_md":null,"page_url":"https://postcutoff.com/m/octo/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":909,"you_url":null,"briefings":null},{"id":"gpt-4o","name":"GPT-4o","org":"OpenAI","family":"GPT-4o","released":"2024-05-13","status":"legacy","type":"multimodal","modality_in":["text","image"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":128000,"max_output":16384,"knowledge_cutoff":"2023-10","pricing":{"input":2.5,"cached_input":1.25,"output":10,"unit":"per 1M tokens (USD), standard tier","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$2.50 in, $10 out per 1M tokens","access":[{"provider":"OpenAI API","model_id":"gpt-4o","endpoint":"https://api.openai.com/v1/chat/completions","docs":"https://developers.openai.com/api/docs/models/gpt-4o"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"gpt-4o","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"},{"provider":"OpenRouter","model_id":"openai/gpt-4o","url":"https://openrouter.ai/openai/gpt-4o"}],"capabilities":[{"name":"Omni model","detail":"Natively multimodal \"o\" model; basis of the gpt-4o audio/realtime/transcribe/TTS variants.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4o"},{"name":"Fine-tunable","detail":"Supports fine-tuning via v1/fine-tuning.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/gpt-4o"}],"entry":null,"notes":"Snapshots gpt-4o-2024-11-20, -2024-08-06, -2024-05-13 (the last deprecated, shutdown Oct 23 2026). chatgpt-4o-latest shut down Feb 17 2026. gpt-4o-mini ($0.15/$0.60) still listed.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/responses \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"gpt-4o\",\"input\":\"Hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/gpt-4o\n- https://developers.openai.com/api/docs/pricing\n- https://developers.openai.com/api/docs/deprecations\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/gpt-4o/","events_after":930,"major_after":256,"historic_after":53,"missing_at_launch":19,"events_since_release":null,"you_url":"https://postcutoff.com/you/gpt-4o/","briefings":{"xs":"https://postcutoff.com/briefings/cutoff-2023-10-xs.md","s":"https://postcutoff.com/briefings/cutoff-2023-10-s.md","m":"https://postcutoff.com/briefings/cutoff-2023-10-m.md","full":"https://postcutoff.com/briefings/cutoff-2023-10.md"}},{"id":"text-embedding-3-large","name":"text-embedding-3-large","org":"OpenAI","family":"text-embedding-3","released":"2024-01-25","status":"current","type":"embedding","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.13,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.13 in per 1M tokens","access":[{"provider":"OpenAI API","model_id":"text-embedding-3-large","endpoint":"https://api.openai.com/v1/embeddings","docs":"https://developers.openai.com/api/docs/models/text-embedding-3-large"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"text-embedding-3-large","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Multilingual embeddings","detail":"Most capable OpenAI embedding model for English and non-English tasks.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/text-embedding-3-large"},{"name":"Shortenable (Matryoshka-style) vectors","detail":"Default 3072 dimensions; the `dimensions` API parameter truncates embeddings while keeping semantic quality. Max input 8192 tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/guides/embeddings"}],"entry":null,"notes":"Output is an embedding vector. Still OpenAI's newest embedding model as of 2026-09.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/embeddings \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"text-embedding-3-large\",\"input\":\"hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/text-embedding-3-large\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/text-embedding-3-large/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":918,"you_url":null,"briefings":null},{"id":"text-embedding-3-small","name":"text-embedding-3-small","org":"OpenAI","family":"text-embedding-3","released":"2024-01-25","status":"current","type":"embedding","modality_in":["text"],"modality_out":["text"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"input":0.02,"unit":"per 1M tokens (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.02 in per 1M tokens","access":[{"provider":"OpenAI API","model_id":"text-embedding-3-small","endpoint":"https://api.openai.com/v1/embeddings","docs":"https://developers.openai.com/api/docs/models/text-embedding-3-small"},{"provider":"Azure OpenAI (Microsoft Foundry)","model_id":"text-embedding-3-small","docs":"https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure"}],"capabilities":[{"name":"Cheap embeddings","detail":"Improved successor to ada-002 at $0.02 per 1M tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/text-embedding-3-small"},{"name":"Shortenable vectors","detail":"Default 1536 dimensions; can be shortened with the `dimensions` parameter. Max input 8192 tokens.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/guides/embeddings"}],"entry":null,"notes":"Output is an embedding vector.","verified":"2026-09-29","body_md":"```bash\ncurl https://api.openai.com/v1/embeddings \\\n -H \"Authorization: Bearer $OPENAI_API_KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\":\"text-embedding-3-small\",\"input\":\"hello\"}'\n```\n\nSources:\n- https://developers.openai.com/api/docs/models/text-embedding-3-small\n- https://developers.openai.com/api/docs/pricing\n- Models overview: https://developers.openai.com/api/docs/models","page_url":"https://postcutoff.com/m/text-embedding-3-small/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":918,"you_url":null,"briefings":null},{"id":"tesla-optimus-ai","name":"Tesla Optimus AI (end-to-end robot neural network)","org":"Tesla","family":"Optimus","released":"2024","status":"preview","type":"robotics","modality_in":["image","video","text"],"modality_out":["action"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Not available (internal only)","url":"https://www.tesla.com/AI"}],"capabilities":[{"name":"Camera-only end-to-end policy on the FSD computer","detail":"Tesla-published Optimus demos (e.g. battery-cell sorting) are described as a single end-to-end neural network running on the robot's onboard FSD computer from camera (and touch) input; Tesla shares the vision/AI stack with FSD.","first":false,"discovered":"launch","source":"https://en.wikipedia.org/wiki/Optimus_(robot)"},{"name":"Offline autonomy on AI5, Grok for conversation","detail":"On the Q1 2026 call (2026-04-22) Musk said the AI5 chip should give Optimus enough local intelligence to keep working without connectivity, while Grok-level conversation needs WiFi/cellular.","first":false,"discovered":"later","source":"https://en.wikipedia.org/wiki/Optimus_(robot)"}],"entry":null,"notes":"Not a product you can call: Tesla has published no model name, architecture, paper, API or weights for the Optimus neural net; this file tracks the robot AI stack. Hardware status (as of 2026-09-29): Optimus V3 / Gen 3 has NOT been unveiled. Tesla missed its Q1 2026 and \"mid-2026\" reveal targets; Musk said on 2026-04-22 it \"will be unveiled closer to production start\" and that Tesla is holding back demos because competitors copy them frame by frame. Tesla's Q1 2026 update says Fremont (former Model S/X line) is being fitted for a 1M-robot/yr first-generation line, with a Giga Texas line targeting 10M/yr long term from 2027. Rumoured V3 specs (22-DoF hands, ~$20-30K price, public sale end-2027) come from secondary sources and are unverified. Sources: Tesla Q1 2026 update and earnings call via https://en.wikipedia.org/wiki/Optimus_(robot) ; https://electrek.co/2026/04/22/tesla-optimus-production-fremont-model-sx-line/ ; https://driveteslacanada.ca/news/tesla-delaying-optimus-v3-reveal-fears-copycats/","verified":null,"body_md":"Tesla's humanoid robot \"brain\". No public access of any kind. Production report (The Information, 2026-09-25; entry 2026-09-25-tesla-optimus-production-problems): several hundred V3-generation units/week by August 2026, mostly for internal testing and data collection; AI still does not generalize beyond trained tasks. Re-check after the Optimus V3 unveil (expected late 2026 per analysts, unconfirmed).","page_url":"https://postcutoff.com/m/tesla-optimus-ai/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":886,"you_url":null,"briefings":null},{"id":"tts-1","name":"TTS-1 / TTS-1 HD","org":"OpenAI","family":"OpenAI TTS","released":"2023-11-06","status":"legacy","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1m_characters":15,"unit":"USD per 1M characters for tts-1; tts-1-hd $30 per 1M characters","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$15 per 1M characters","access":[{"provider":"OpenAI API","model_id":"tts-1","endpoint":"https://api.openai.com/v1/audio/speech","docs":"https://developers.openai.com/api/docs/models/tts-1"},{"provider":"OpenAI API","model_id":"tts-1-hd","endpoint":"https://api.openai.com/v1/audio/speech","docs":"https://developers.openai.com/api/docs/models/tts-1-hd"}],"capabilities":[{"name":"Low-latency preset-voice TTS","detail":"tts-1 optimised for real-time synthesis; tts-1-hd for higher quality at twice the price.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/tts-1"}],"entry":null,"notes":"Not deprecated as of 2026-09-29 but no longer shown in the models overview; gpt-4o-mini-tts is the current, instruction-steerable replacement. Release date = OpenAI DevDay 2023 (from memory, not re-verified today).","verified":"2026-09-29","body_md":"Sources:\n- https://developers.openai.com/api/docs/models/tts-1\n- https://developers.openai.com/api/docs/pricing","page_url":"https://postcutoff.com/m/tts-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":929,"you_url":null,"briefings":null},{"id":"whisper-large-v3","name":"Whisper large-v3 / large-v3-turbo (open weights)","org":"OpenAI","family":"Whisper","released":"2023-11-06","status":"legacy","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Hugging Face","model_id":"openai/whisper-large-v3","url":"https://huggingface.co/openai/whisper-large-v3"},{"provider":"Hugging Face","model_id":"openai/whisper-large-v3-turbo","url":"https://huggingface.co/openai/whisper-large-v3-turbo"},{"provider":"GitHub","url":"https://github.com/openai/whisper"},{"provider":"Groq","model_id":"whisper-large-v3-turbo","docs":"https://console.groq.com/docs/speech-to-text"},{"provider":"Deepgram (hosted)","model_id":"whisper-large","docs":"https://developers.deepgram.com/docs/models-languages-overview"}],"capabilities":[{"name":"Robust multilingual ASR + translation to English","detail":"99 languages; timestamps; zero-shot speech translation into English; the de facto open ASR baseline.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/whisper-large-v3"},{"name":"Turbo: 4-layer decoder","detail":"large-v3-turbo (Oct 2024) prunes the decoder from 32 to 4 layers (809M vs 1.55B params) for much faster decoding with minor quality loss; not trained for translation.","first":false,"discovered":"launch","source":"https://huggingface.co/openai/whisper-large-v3-turbo"}],"entry":null,"notes":"Status legacy: still widely deployed, but surpassed on the Open ASR Leaderboard by NVIDIA Canary/Parakeet, and OpenAI's API now points to gpt-transcribe (whisper-1 API shutdown 2027-02-26). Known to hallucinate text on silence/noise. Dates from OpenAI releases (large-v3 at DevDay 2023-11-06; turbo 2024-10-01), not re-checked today.","verified":"2026-09-29","body_md":"Sources: https://huggingface.co/openai/whisper-large-v3-turbo , https://github.com/openai/whisper","page_url":"https://postcutoff.com/m/whisper-large-v3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":929,"you_url":null,"briefings":null},{"id":"elevenlabs-multilingual-v2","name":"Eleven Multilingual v2","org":"ElevenLabs","family":"Eleven v2","released":"2023-08-22","status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.08,"unit":"per 1K characters (API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.08 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_multilingual_v2","endpoint":"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/models"},{"provider":"Web app","url":"https://elevenlabs.io/app"}],"capabilities":[{"name":"Stable long-form multilingual TTS","detail":"'Lifelike model with rich emotional expression', 29 languages, 10,000 chars/request; keeps a voice's characteristics across languages. Launched as ElevenLabs exited beta.","first":false,"discovered":"launch","source":"https://elevenlabs.io/blog/elevenlabs-comes-out-of-beta-and-releases-eleven-multilingual-v2-a-foundational-ai-speech-model-for-nearly-30-languages"}],"entry":null,"notes":"Launched 2023-08-22 (press date). Still current and the long-standing default for narration; supports style/speed settings and PVC. Superseded in expressiveness by v3/v4.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api","page_url":"https://postcutoff.com/m/elevenlabs-multilingual-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":932,"you_url":null,"briefings":null},{"id":"whisper-1","name":"Whisper (whisper-1 API)","org":"OpenAI","family":"Whisper","released":"2023-03-01","status":"deprecated","type":"audio/speech","modality_in":["audio"],"modality_out":["text"],"open_weights":true,"model_license":"mit","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.006,"unit":"per minute of audio (USD)","source":"https://developers.openai.com/api/docs/pricing"},"price_line":"$0.006 per minute","access":[{"provider":"OpenAI API","model_id":"whisper-1","endpoint":"https://api.openai.com/v1/audio/transcriptions (also /v1/audio/translations)","docs":"https://developers.openai.com/api/docs/models/whisper-1"},{"provider":"GitHub (open weights)","url":"https://github.com/openai/whisper"},{"provider":"Hugging Face","url":"https://huggingface.co/openai/whisper-large-v3"}],"capabilities":[{"name":"Multilingual speech recognition, translation and language ID","detail":"General-purpose ASR trained on a large diverse audio dataset; transcribes many languages and translates speech into English.","first":false,"discovered":"launch","source":"https://developers.openai.com/api/docs/models/whisper-1"}],"entry":null,"notes":"API model deprecated 2026-08-26, shutdown 2027-02-26; replacements gpt-transcribe / gpt-live-transcribe. The open-source Whisper checkpoints (MIT, first released Sept 2022) remain downloadable and widely self-hosted; the API's whisper-1 has no snapshot versions. API launch date (March 2023, with the ChatGPT API) is from memory, not re-verified today.","verified":"2026-09-29","body_md":"Sources:\n- https://developers.openai.com/api/docs/models/whisper-1\n- https://developers.openai.com/api/docs/deprecations\n- https://github.com/openai/whisper","page_url":"https://postcutoff.com/m/whisper-1/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":943,"you_url":null,"briefings":null},{"id":"elevenlabs-voice-changer-sts-v2","name":"Eleven Multilingual STS v2 (Voice Changer)","org":"ElevenLabs","family":"Eleven v2","released":null,"status":"current","type":"audio/speech","modality_in":["audio"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_minute":0.12,"unit":"per minute of audio (Voice Changer API)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.12 per minute","access":[{"provider":"ElevenLabs API","model_id":"eleven_multilingual_sts_v2","endpoint":"https://api.elevenlabs.io/v1/speech-to-speech/{voice_id}","docs":"https://elevenlabs.io/docs/overview/capabilities/voice-changer"},{"provider":"ElevenLabs API (English only)","model_id":"eleven_english_sts_v2","endpoint":"https://api.elevenlabs.io/v1/speech-to-speech/{voice_id}"},{"provider":"Web app","url":"https://elevenlabs.io/voice-changer"}],"capabilities":[{"name":"Speech-to-speech voice conversion","detail":"Converts a recording into another voice while keeping the original delivery (timing, emotion); multilingual model covers 29 languages.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"}],"entry":null,"notes":"Release date not verified. Voice Isolator is also $0.12/min.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/pricing/api","page_url":"https://postcutoff.com/m/elevenlabs-voice-changer-sts-v2/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":null,"you_url":null,"briefings":null},{"id":"elevenlabs-turbo-v2-5","name":"Eleven Turbo v2.5 / Turbo v2","org":"ElevenLabs","family":"Eleven Turbo","released":null,"status":"deprecated","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":{"per_1k_characters":0.04,"unit":"per 1K characters (API, Flash/Turbo tier)","source":"https://elevenlabs.io/pricing/api"},"price_line":"$0.04 per 1K characters","access":[{"provider":"ElevenLabs API","model_id":"eleven_turbo_v2_5","docs":"https://elevenlabs.io/docs/models"},{"provider":"ElevenLabs API (English only)","model_id":"eleven_turbo_v2","docs":"https://elevenlabs.io/docs/models"}],"capabilities":[],"entry":null,"notes":"Marked deprecated on the models page: 'First generation low-latency model (outclassed by Flash)'. Turbo v2.5: 32 languages; Turbo v2: English only. Migrate to eleven_flash_v2_5 or eleven_v4_turbo. Release dates (2024) not re-verified; shutdown date not stated.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models","page_url":"https://postcutoff.com/m/elevenlabs-turbo-v2-5/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":null,"you_url":null,"briefings":null},{"id":"elevenlabs-voice-design-v3","name":"Eleven Voice Design v3 (Text to Voice)","org":"ElevenLabs","family":"Eleven v3","released":null,"status":"current","type":"audio/speech","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"ElevenLabs API","model_id":"eleven_ttv_v3","endpoint":"https://api.elevenlabs.io/v1/text-to-voice/design","docs":"https://elevenlabs.io/docs/models"},{"provider":"ElevenLabs API (older)","model_id":"eleven_multilingual_ttv_v2","endpoint":"https://api.elevenlabs.io/v1/text-to-voice/design"},{"provider":"Web app","url":"https://elevenlabs.io/voice-design"}],"capabilities":[{"name":"Design a voice from a text description","detail":"Generates new synthetic voices from a prompt; eleven_ttv_v3 covers 70+ languages, eleven_multilingual_ttv_v2 29.","first":false,"discovered":"launch","source":"https://elevenlabs.io/docs/models"}],"entry":null,"notes":"Pricing not listed on the API pricing page (billed in credits). Release date not verified. ElevenLabs warns Voice Design voices may not perform as well on Eleven v4 as on earlier models.","verified":"2026-09-29","body_md":"Sources: https://elevenlabs.io/docs/models , https://elevenlabs.io/docs/overview/capabilities/text-to-speech/eleven-v4","page_url":"https://postcutoff.com/m/elevenlabs-voice-design-v3/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":null,"you_url":null,"briefings":null},{"id":"lyria-realtime","name":"Lyria RealTime","org":"Google DeepMind","family":"Lyria","released":null,"status":"preview","type":"music","modality_in":["text"],"modality_out":["audio"],"open_weights":false,"model_license":"proprietary","context_window":null,"max_output":null,"knowledge_cutoff":null,"pricing":null,"price_line":"Price not published","access":[{"provider":"Gemini API (Live music, WebSocket)","model_id":"models/lyria-realtime-exp","docs":"https://ai.google.dev/gemini-api/docs/realtime-music-generation"}],"capabilities":[{"name":"Interactive streaming music generation","detail":"Persistent bidirectional WebSocket session that continuously streams 48 kHz stereo 16-bit PCM; steer live with weighted text prompts and play/pause/stop/reset controls.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/realtime-music-generation"},{"name":"Live musical parameters","detail":"Adjust guidance (0-6), BPM (60-200), density, brightness, scale (12 key pairs) and mute bass/drums on the fly.","first":false,"discovered":"launch","source":"https://ai.google.dev/gemini-api/docs/realtime-music-generation"}],"entry":null,"notes":"Experimental model (status 'Experimental' on the Gemini API models page; no shutdown date announced). Instrumental only; output is SynthID-watermarked. No price listed on the Gemini API pricing page as of 2026-09-29. Release date not re-verified here (it first appeared in 2025 as an experimental model).","verified":"2026-09-29","body_md":"Real-time, steerable instrumental music stream for apps, installations and live performance. Use `client.aio.live.music.connect(model=\"models/lyria-realtime-exp\")` in the Google Gen AI SDK.\n\nSources: [Lyria RealTime guide](https://ai.google.dev/gemini-api/docs/realtime-music-generation), [models](https://ai.google.dev/gemini-api/docs/models), [deprecations](https://ai.google.dev/gemini-api/docs/deprecations).","page_url":"https://postcutoff.com/m/lyria-realtime/","events_after":null,"major_after":null,"historic_after":null,"missing_at_launch":null,"events_since_release":null,"you_url":null,"briefings":null}]}