GPT-Realtime-2
Launched 2026-05-07 with gpt-realtime-translate and gpt-realtime-whisper (changelog). Superseded two months later by gpt-realtime-2.1 (2026-07-06) at identical prices, but still listed and not deprecated. Realtime endpoint only; function calling and prompt caching. Official launch post (openai.com) returned 403 to our fetcher, so benchmark claims were not read directly; secondary sources quote OpenAI: +15.2% Big Bench Audio vs gpt-realtime-1.5 (high effort), +13.8% Audio MultiChallenge instruction following (xhigh); one blog reports 96.6% absolute Big Bench Audio at xhigh (unconfirmed).
- Context window
- 128,000 tokens
- Max output
- 32,000 tokens
- Knowledge cutoff
- 2024-09
- Input
- text, audio, image
- Output
- text, audio
- License
- proprietary
- Pricing
- text input: $4 · text output: $24 · audio input: $32 · audio output: $64 · cached input: $0.4 · image input: $5 (per 1M tokens (USD); cached audio/text input $0.40, cached image $0.50) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| OpenAI API | gpt-realtime-2 | wss://api.openai.com/v1/realtime | docs |
Notable capabilities (2)
- Reasoning speech-to-speech model: First OpenAI realtime voice model with configurable reasoning effort (press: 'GPT-5-class' reasoning); higher effort adds latency and tokens. source
- 128K-token realtime context: Context grew from 32K (gpt-realtime-1.5) to 128K tokens, with 32K max output, for long voice-agent sessions. source
OpenAI's first reasoning realtime voice model. For new builds prefer gpt-realtime-2.1.
Sources:
Timeline entry
- OpenAI releases GPT-Realtime-2 (reasoning voice), GPT-Realtime-Translate and GPT-Realtime-Whisper ★★★
On 2026-05-07 OpenAI added three streaming audio models to its Realtime API: gpt-realtime-2, its first speech-to-speech model with configurable reasoning effort and a 128K context; gpt-realtime-translate for live speech-to-speech interpretation (70+ input, 13 output languages, $0.034/min); and…
Other OpenAI models
GPT-6 Luna · GPT-6 Sol · GPT Image 2.5 Flare · GPT Image 2.5 Sunburst · GPT-6 Astra · GPT-Live-Transcribe · GPT-Transcribe · GPT-5.6 Terra · GPT-Live 1 · GPT-Realtime-2.1 · GPT-Realtime-Translate · GPT-Rosalind · GPT-Audio-1.5 (and gpt-audio / gpt-audio-mini) · GPT-5.3-Codex · gpt-oss-120b · gpt-oss-20b · GPT-4o mini TTS · text-embedding-3-large · text-embedding-3-small · GPT-5.6 Luna · GPT-5.6 Sol · GPT-Realtime-Whisper · GPT-5.5 Pro · GPT-5.5 · GPT Image 2 · GPT-5.4 · GPT-Realtime-1.5 · GPT-4.1 · GPT-4o · TTS-1 / TTS-1 HD · Whisper large-v3 / large-v3-turbo (open weights) · GPT-Realtime and GPT-Realtime mini · o3 · GPT-4o Transcribe / Mini Transcribe / Transcribe Diarize · Whisper (whisper-1 API) · Sora 2