Inworld Realtime TTS-2 reaches GA with audio-aware, prompt-directed speech
Inworld AI made Realtime TTS-2 (`inworld-tts-2`) and TTS-2 Flash generally available on 2026-08-31 after a research preview on 2026-05-05. TTS-2 conditions on the actual audio of earlier turns, so it can pick up a user's tone and pacing. It takes plain-English voice direction and keeps one voice identity across 100+ languages, at $25 (Flash $15) per 1M characters pay-as-you-go.
Key facts
- Model id inworld-tts-2; endpoint POST https://api.inworld.ai/tts/v1/voice
- TTS-2 median TTFA <200 ms; Flash ~20 ms TTFB (docs)
- Voice cloning from 5-15 s; voice design from text; STABLE/BALANCED/CREATIVE modes
- Artificial Analysis 29 Sept 2026: #5 (Elo 1244); Inworld's earlier TTS 1.5 had been #1
- TTS-1..1.5 discontinued 2026-06-15; Inworld also offers migration from shut-down PlayHT
What happened
Inworld promoted TTS-2 from research preview to GA and added a Flash variant for latency- and cost-sensitive agents.
Why it matters
TTS-2 closes the loop between listening and speaking in a cascaded voice stack: the TTS hears the user, not just the transcript. The price is also well under ElevenLabs' list price.
Changelog
- 2026-09-29: created
Models
- Inworld Realtime TTS-2 / TTS-2 Flash Inworld AI · current
Related events
- Cartesia Sonic-3.6 goes GA and tops the Artificial Analysis Speech Arena ★★★
- ElevenLabs launches Eleven v4 and Eleven v4 Turbo, #1 on Artificial Analysis TTS arena ★★★★
- Meta acquires voice-AI startup PlayAI (PlayHT); the product is later shut down ★★
- Alibaba's Qwen-Audio-3.0-TTS takes #1 on the Artificial Analysis text-to-speech leaderboard ★★★
- Deepgram launches Flux TTS and passes $100M ARR ★★★
Sources (4)
- officialInworld: Realtime TTS-2
- docsInworld docs: TTS models
- officialInworld pricing
- pressMarkTechPost: preview launch (2026-05-05)
id: 2026-08-31-inworld-realtime-tts-2 · updated 2026-09-29 · open in the interactive timeline