Grok Voice Transcribe 2.0
Launched 2026-09-18 as drop-in upgrade of the Grok STT API (first released 2026-04-17 with grok-voice-transcribe-1.0, which can be pinned but will be deprecated). Up to 500 MB files; WAV/MP3/OGG/Opus/FLAC/AAC/MP4/M4A/MKV plus raw PCM/mu-law/A-law at 8-48 kHz; Smart Turn end-of-turn detection, VAD, inverse text normalization, filler removal, mid-recording language switching. Docs list ~25 languages for formatting.
- Input
- audio
- Output
- text
- License
- proprietary
- Pricing
- per hour batch: $0.1 · per hour streaming: $0.2 (USD per hour of audio (REST batch / WebSocket streaming); diarization, timestamps and key-term biasing included) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| xAI API (REST) | grok-voice-transcribe-2.0 | https://api.x.ai/v1/stt | docs |
| xAI API (WebSocket streaming) | grok-voice-transcribe-2.0 | wss://api.x.ai/v1/stt | docs |
Notable capabilities (2)
- Top streaming STT accuracy (claimed): xAI says it ranks #1 for accuracy among 32 streaming models on the Artificial Analysis leaderboard; multilingual short-phrase WER 20.6% -> 6.8% vs v1.0 ('2x as accurate'). source
- Very low price with diarization included: $0.10/hr batch and $0.20/hr streaming, with speaker diarization, word timestamps, up to 8-channel multichannel and 100 key terms per request at no extra cost. source
curl https://api.x.ai/v1/stt -H "Authorization: Bearer $XAI_API_KEY" \
-F model=grok-voice-transcribe-2.0 -F file=@call.wav -F diarize=true
Sources: https://x.ai/news/grok-voice-transcribe-2 · https://docs.x.ai/developers/model-capabilities/audio/speech-to-text · https://x.ai/news/grok-stt-and-tts-apis
Other xAI models
Grok 4.7 · Grok Imagine Image 2.0 · Grok Voice Think Fast 2.0 · Grok Imagine Video 1.5 · Grok Build 0.1 · Grok Text to Speech (Grok TTS API) · Grok 4.3 · Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)