Inworld Realtime TTS-2 / TTS-2 Flash
Research preview 2026-05-05, GA 2026-08-31. Inworld claimed #1 on Artificial Analysis Speech Arena; on 2026-09-29 AA shows it #5 (Elo 1244) behind Eleven v4, Sonic 3.6, Gemini 3.8 Flash TTS, Qwen-Audio-3.0-TTS-Plus. Docs say 200+ languages vs 100+ in blog. TTS-1..1.5 discontinued 2026-06-15 (auto-routed). Flash model id not verified. Max 2,000 chars/request.
- Input
- text, audio
- Output
- audio
- License
- proprietary
- Pricing
- per 1m characters: $25 (USD per 1M characters PAYG for TTS-2 ($15 for TTS-2 Flash); plan rates down to $12.50/$7, enterprise from $5) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Inworld API | inworld-tts-2 | https://api.inworld.ai/tts/v1/voice | docs |
| Cloudflare Workers AI | — | developers.cloudflare.com/ai/models/inworld/tts-2/ | — |
Notable capabilities (3)
- Closed-loop, audio-aware delivery: Conditions on the actual audio of prior turns (user tone, pacing, emotion), not just transcripts, and takes plain-English voice direction; delivery modes STABLE/BALANCED/CREATIVE. source
- Cross-lingual identity in 100+ languages: One voice holds identity while switching language on the fly; cloning from 5-15 s reference or voice design from a text description. source
- Flash variant ~20 ms TTFB: TTS-2 Flash: ~20 ms TTFB, ~5x faster than inworld-tts-2 (docs); TTS-2 median TTFA under 200 ms. source
Sources: https://inworld.ai/blog/realtime-tts-2 , https://docs.inworld.ai/tts/tts-models , https://inworld.ai/pricing
Timeline entry
- Inworld Realtime TTS-2 reaches GA with audio-aware, prompt-directed speech ★★★
Inworld AI made Realtime TTS-2 (`inworld-tts-2`) and TTS-2 Flash generally available on 2026-08-31 after a research preview on 2026-05-05. TTS-2 conditions on the actual audio of earlier turns, so it can pick up a user's tone and pacing. It takes plain-English voice direction and keeps one voice…