Soniox v5 (Async and Real-Time STT)
stt-async-v5 released 2026-06-11, stt-rt-v5 on 2026-06-16. The v4 ids (stt-async-v4 from 2026-01-29, stt-rt-v4 from 2026-02-05) were retired 2026-06-30 and are now aliases routing to v5. Launch posts give no WER numbers; Soniox publishes its own comparisons at soniox.com/benchmarks (vendor-run). Sibling TTS: soniox-tts-v2.
- Input
- audio, text
- Output
- text
- License
- proprietary
- Pricing
- input: $1.5 · output: $3.5 (USD per 1M tokens for async (audio in $1.50, text in/out $3.50; ~$0.10 per audio hour). Real-time: $2.00 audio in, $4.00 text in/out (~$0.12/hour). 1 hour of audio ≈ 30,000 input tokens.) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Soniox API (async / file) | stt-async-v5 | — | docs |
| Soniox API (real-time streaming) | stt-rt-v5 | — | docs |
| Web app | — | soniox.com | — |
Notable capabilities (2)
- One multilingual model for 60+ languages with speaker separation: Soniox claims native-speaker accuracy across 60+ languages in a single model, re-engineered speaker diarization, spoken-language ID, context injection and precise alphanumerics (IDs, emails, codes). source
- Real-time translation and semantic endpointing: stt-rt-v5 transcribes and translates live across ~3,600 language pairs, with a tunable `endpoint_sensitivity` semantic endpointing parameter for voice agents. source
Soniox's speech-to-text generation for 2026: a single multilingual model family for files (async) and streaming (real-time).
Sources: https://soniox.com/blog/soniox-v5-async , https://soniox.com/blog/soniox-v5-real-time , https://soniox.com/docs/stt/models , https://x.com/soniox_ai/status/2065083564027257182