Grok Text to Speech (Grok TTS API)
Launched with the Grok STT API on 2026-04-17 (some press reports an earlier developer opening in March 2026). No separate model id is documented; the endpoint selects the model. 60,000 characters per REST request; ~20 languages plus auto-detect; MP3/WAV/PCM/mu-law/A-law at 8-48 kHz; voice list via GET /v1/tts/voices (Ara, Eve, Leo, Rex, Sal and many more).
- Input
- text
- Output
- audio
- License
- proprietary
- Pricing
- per million characters: $15 (USD per 1M characters) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| xAI API (REST) | — | https://api.x.ai/v1/tts | docs |
| xAI API (WebSocket streaming) | — | wss://api.x.ai/v1/tts | docs |
Notable capabilities (2)
- Inline speech tags: Inline tags ([pause], [laugh], [sigh], [cry], [gasp], ...) and wrapping tags (<whisper>, <soft>, <loud>, <slow>, <fast>, <sing>) control delivery. source
- Custom (cloned) voices (found after launch): Clone a voice from a short reference clip via the Custom Voices API; the voice_id works like built-in voices in TTS and the Voice Agent API. source
Call POST https://api.x.ai/v1/tts with your text and a voice (see the docs for the exact request schema).
Sources: https://docs.x.ai/developers/model-capabilities/audio/text-to-speech · https://x.ai/news/grok-stt-and-tts-apis · https://docs.x.ai/developers/pricing
Other xAI models
Grok 4.7 · Grok Voice Transcribe 2.0 · Grok Imagine Image 2.0 · Grok Voice Think Fast 2.0 · Grok Imagine Video 1.5 · Grok Build 0.1 · Grok 4.3 · Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)