Microsoft AI launches MAI-Transcribe-2-Streaming (real-time ASR) and MAI-Voice-2.1 / 2.1-Flash
On Oct 1, 2026 Microsoft AI released MAI-Transcribe-2-Streaming, its first streaming speech-recognition model (first hypotheses in just over 100 ms, 60 languages, $0.54/hour through year-end), and two text-to-speech models, MAI-Voice-2.1 ($22 per 1M characters) and MAI-Voice-2.1-Flash (150 ms end-to-end, $15 per 1M characters). It extends Microsoft's in-house voice stack beyond OpenAI models.
Key facts
- MAI-Transcribe-2-Streaming: first hypotheses in just over 100 ms; Microsoft claims words appear 2x faster than the closest competitor and #1 accuracy for final and partial transcripts on Artificial Analysis
- MAI-Transcribe-2-Streaming: 60 languages with automatic language detection; $0.54 per hour of audio through end of 2026
- MAI-Voice-2.1: 23 languages / 26 locales, one voice keeps its native accent across languages, voice cloning with consent guardrails; $22 per 1M characters
- MAI-Voice-2.1-Flash: 150 ms end-to-end latency, 55% faster inference; $15 per 1M characters
- Availability: Microsoft Foundry, MAI Playground, OpenRouter, Vercel, Azure Voice Live; LiveKit coming soon
What happened
Microsoft AI (Mustafa Suleyman's group) shipped a low-latency streaming version of MAI-Transcribe-2 plus an updated voice-generation pair. All three were available on launch day. Benchmark claims are Microsoft's own, citing Artificial Analysis.
Why it matters
Real-time transcription and TTS are the building blocks of voice agents. Microsoft now offers its own models for both at low prices, competing with OpenAI, ElevenLabs and Deepgram.
Changelog
- 2026-10-02: created
Models
- MAI-Transcribe-2-Streaming Microsoft · current
- MAI-Voice-2.1 / MAI-Voice-2.1-Flash Microsoft · current
Related posts (1)
- Microsoft AI original ↗ Microsoft AI @microsoftai · x · 2026-10-01
Cited as a source by: 2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1
Related events
- Microsoft launches MAI-Transcribe-2, claiming the most accurate and cheapest speech recognition at $0.10/hour ★★★
- Microsoft launches seven in-house MAI models at Build 2026, led by MAI-Thinking-1 ★★★★
- Microsoft makes its own Azure Realtime speech-to-speech model generally available in the Voice Live API ★★
Sources (2)
- officialMicrosoft AI: Our first streaming transcription model
- official@MicrosoftAI on X: all three models available today
id: 2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1 · updated 2026-10-02 · open in the interactive timeline