Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Microsoft AI launches MAI-Transcribe-2-Streaming…

Microsoft AI launches MAI-Transcribe-2-Streaming (real-time ASR) and MAI-Voice-2.1 / 2.1-Flash

★★after cutoffmodel-releaseMicrosoftconfidence: high

On Oct 1, 2026 Microsoft AI released MAI-Transcribe-2-Streaming, its first streaming speech-recognition model (first hypotheses in just over 100 ms, 60 languages, $0.54/hour through year-end), and two text-to-speech models, MAI-Voice-2.1 ($22 per 1M characters) and MAI-Voice-2.1-Flash (150 ms end-to-end, $15 per 1M characters). It extends Microsoft's in-house voice stack beyond OpenAI models.

Key facts

What happened

Microsoft AI (Mustafa Suleyman's group) shipped a low-latency streaming version of MAI-Transcribe-2 plus an updated voice-generation pair. All three were available on launch day. Benchmark claims are Microsoft's own, citing Artificial Analysis.

Why it matters

Real-time transcription and TTS are the building blocks of voice agents. Microsoft now offers its own models for both at low prices, competing with OpenAI, ElevenLabs and Deepgram.

Changelog

  • 2026-10-02: created

Models

Related posts (1)

Related events

  1. Microsoft launches MAI-Transcribe-2, claiming the most accurate and cheapest speech recognition at $0.10/hour ★★★
  2. Microsoft launches seven in-house MAI models at Build 2026, led by MAI-Thinking-1 ★★★★
  3. Microsoft makes its own Azure Realtime speech-to-speech model generally available in the Voice Live API ★★

Sources (2)

id: 2026-10-01-microsoft-mai-transcribe-2-streaming-voice-2-1 · updated 2026-10-02 · open in the interactive timeline