Post-Cutoff.com
  1. Home
  2. Models
  3. Grok Voice Think Fast 2.0

Grok Voice Think Fast 2.0

xAIcurrentaudio/speechGrok Voice

Released 2026-07-29; grok-voice-latest switched to it on 2026-08-05. Predecessor grok-voice-think-fast-1.0 can still be pinned. 20+ languages; audio PCM (8-48 kHz), Opus 24 kHz, G.711 mu-law/A-law; server VAD, session resumption (30 min), custom cloned voices. xAI says Starlink A/B tests raised sales conversion and support containment. Benchmarks are xAI-reported.

Input
audio, text
Output
audio, text
License
proprietary
Pricing
per minute: $0.08 (USD per minute of audio ($4.80/hr), plus $0.004 per text input (as listed on the xAI pricing page)) source
Verified
2026-09-29

How to call it

ProviderModel idEndpoint / URLDocs
xAI API (Voice Agent / speech-to-speech, WebSocket)grok-voice-think-fast-2.0wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0docs
xAI API (alias)grok-voice-latest—docs
Web app—grok.com—

Notable capabilities (4)

xAI's realtime voice-agent model (also used in the Grok app, Tesla vehicles and Starlink support).

# OpenAI Realtime-compatible WebSocket
import websockets, os
url = "wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0"
ws = await websockets.connect(url, additional_headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"})

Sources: https://x.ai/news/grok-voice-think-fast-2 · https://docs.x.ai/developers/model-capabilities/audio/voice-agent · https://docs.x.ai/developers/pricing

Timeline entry

  1. xAI releases Grok Voice Think Fast 2.0 speech-to-speech model for voice agents ★★★

    On 2026-07-29 xAI (SpaceXAI) released Grok Voice Think Fast 2.0, a reasoning speech-to-speech model for its OpenAI-Realtime-compatible Voice Agent API at $0.08/min, scoring 82.9 on the Artificial Analysis Speech-to-Speech Quality Index and cutting time to first audio to 0.70 s.

Other xAI models

Grok 4.7 · Grok Voice Transcribe 2.0 · Grok Imagine Image 2.0 · Grok Imagine Video 1.5 · Grok Build 0.1 · Grok Text to Speech (Grok TTS API) · Grok 4.3 · Grok 4.20 (Reasoning / Non-reasoning / Multi-Agent)