OpenAI releases GPT-Realtime-2 (reasoning voice), GPT-Realtime-Translate and GPT-Realtime-Whisper
On 2026-05-07 OpenAI added three streaming audio models to its Realtime API: gpt-realtime-2, its first speech-to-speech model with configurable reasoning effort and a 128K context; gpt-realtime-translate for live speech-to-speech interpretation (70+ input, 13 output languages, $0.034/min); and gpt-realtime-whisper for streaming transcription ($0.017/min).
Key facts
- gpt-realtime-2: text $4 / $24, audio $32 / $64 per 1M tokens; 128K context (up from 32K), 32K max output
- gpt-realtime-translate: v1/realtime/translations endpoint, 70+ input and 13 output languages (press), $0.034 per minute
- gpt-realtime-whisper: streaming speech-to-text, tunable latency, $0.017 per minute
- Benchmarks (OpenAI launch post, quoted by secondary sources; post itself 403 to our tools): gpt-realtime-2 (high) +15.2% on Big Bench Audio vs gpt-realtime-1.5; (xhigh) +13.8% on Audio MultiChallenge instruction following. One blog gives 96.6% absolute on Big Bench Audio at xhigh (unconfirmed)
- Superseded by gpt-realtime-2.1 on 2026-07-06 and, for transcription, gpt-live-transcribe on 2026-07-28
What happened
OpenAI shipped three Realtime API models the same day. GPT-Realtime-2 brings adjustable reasoning to speech-to-speech voice agents (press described it as GPT-5-class reasoning) and quadruples the context to 128K tokens. GPT-Realtime-Translate is a dedicated simultaneous-interpretation model billed per minute. GPT-Realtime-Whisper streams transcripts from live audio.
Why it matters
Reasoning moved into the low-latency voice loop instead of being bolted on via a separate text model, and live translation became a standalone API product, a month before Google's Gemini 3.5 Live Translate.
The official post (openai.com) could not be fetched by our tools; language counts come from press coverage.
Changelog
- 2026-09-29: created
- 2026-09-29: added benchmark deltas (Big Bench Audio, Audio MultiChallenge) from secondary quotes of the 403-blocked launch post, plus OpenAI community announcement link
Models
- GPT-Realtime-2 OpenAI · current
- GPT-Realtime-Translate OpenAI · current
- GPT-Realtime-Whisper OpenAI · legacy
Related events
- OpenAI launches GPT-Live, full-duplex voice models replacing ChatGPT's Advanced Voice Mode ★★★★
- Google launches Gemini 3.5 Live Translate, voice-preserving real-time speech translation in 70+ languages ★★★
- ByteDance Seed LiveInterpret 2.0: end-to-end Chinese-English simultaneous interpretation in your own voice, ~3 s behind ★★★
- Microsoft makes its own Azure Realtime speech-to-speech model generally available in the Voice Live API ★★
- OpenAI releases GPT-Transcribe and GPT-Live-Transcribe, then deprecates Whisper API ★★★
Sources (7)
- officialOpenAI - Advancing voice intelligence with new models in the API
- docsOpenAI API changelog
- docsgpt-realtime-2 model page
- docsgpt-realtime-translate model page
- officialOpenAI Developer Community - New Realtime Voice Models in the API
- pressBuild Fast with AI - GPT-Realtime-2 benchmarks (secondary)
- pressgHacks - OpenAI releases three new realtime voice models
id: 2026-05-07-openai-gpt-realtime-2-translate-whisper · updated 2026-09-29 · open in the interactive timeline