OpenAI launches GPT-4o, a natively multimodal 'omni' model
GPT-4o reasoned natively across text, audio and vision in real time, responding to speech in as little as 232 ms, and brought GPT-4-level intelligence to free ChatGPT users.
Key facts
- Announced 13 May 2024
- Audio response latency as low as 232 ms, ~320 ms on average
- Single end-to-end model for text, vision and audio
- Available to free ChatGPT users; half the API price of GPT-4 Turbo
- Advanced Voice Mode rolled out later in 2024
What happened
OpenAI demoed a conversational assistant that could hear, see and speak with human-like latency and emotional expressiveness.
Why it matters
Made natural voice interaction with AI mainstream and set the standard for omni-modal assistants.
Changelog
- 2026-09-29: created
Related events
- OpenAI releases GPT-4 ★★★★★
- OpenAI o1: reasoning models trained with reinforcement learning ★★★★★
- OpenAI launches GPT-Live, full-duplex voice models replacing ChatGPT's Advanced Voice Mode ★★★★
Sources (2)
id: 2024-05-13-gpt-4o · updated 2026-09-29 · open in the interactive timeline