SpaceXAI Ships Grok Voice Think Fast 2.0 With Top Speech-to-Speech Scores
On July 29, 2026, SpaceXAI released Grok Voice Think Fast 2.0—its next-generation speech-to-speech voice model—claiming 82.9% on Artificial Analysis’s Speech-to-Speech Quality Index, 0.70s time-to-first-audio, $0.08/min pricing, and a default swap for grok-voice-latest on August 5.
TLDR
SpaceXAI on July 29, 2026 announced Grok Voice Think Fast 2.0, its most capable speech-to-speech voice model. The company primary cites Artificial Analysis leadership on speech-to-speech quality (82.9% vs 79.1% for GPT-Realtime-2.1 High), 0.70s time to first audio (down from 1.25s on 1.0), roughly 0.4× the reasoning tokens of its predecessor in production, and $0.08 per minute of audio. grok-voice-latest migrates to 2.0 on August 5, 2026 unless pinned.
What shipped
| Item | Detail (x.ai primary) |
|---|---|
| Model | Grok Voice Think Fast 2.0 |
| Focus | Speech reasoning, transcription accuracy, conversational tool use |
| AA Speech-to-Speech Quality | 82.9% (1.0: 75.7%; GPT-Realtime-2.1 High: 79.1%) |
| tau-Voice agentic | 56.5% vs 45.7% for GPT-Realtime-2.1 High (lab table) |
| Latency | 0.70s time to first audio (1.0: 1.25s) |
| Pricing | $0.08 / min of audio |
| Default cutover | grok-voice-latest → 2.0 on Aug 5, 2026 (pin grok-voice-think-fast-1.0 to stay) |
| Production claim | A/B on Starlink support/sales lines: higher conversion and containment |
SpaceXAI positions 2.0 as reasoning while speaking—tool calls often fire before the first sentence finishes—without adding turn-taking latency. Transcription claims include 1.5–2.0× lower word-error relative to Deepgram Nova 3 and ElevenLabs Scribe v2 on internal multilingual short-phrase tests, with larger gaps under noise/telephony compression.
Independent rankings
| Source | Metric | Snapshot |
|---|---|---|
| Artificial Analysis Speech-to-Speech | Quality Index 82.9% | Cited on x.ai launch post as leading GPT-Realtime-2.1 High (79.1%) and Gemini 3.1 Flash High (69.5%) |
| AA / lab tau-Voice | Agentic voice 56.5% | Lab table vs 45.7% (GPT-Realtime-2.1 High) |
Product-line placement
Distinct from Grok 4.5 (text/coding flagship), Build Mode, and Imagine media models. Continues the Voice Agent API / Think Fast line launched earlier in 2026 (1.0). Competes with OpenAI GPT-Live / Realtime and Google Gemini voice stacks on full-duplex speech agents.
Why this story matters
Voice is becoming a primary enterprise surface for support, sales, and agent workflows—not a demo. Shipping a measurable step-change on third-party speech-to-speech boards and a default API migration date forces teams to re-evaluate cost, latency, and pin strategy before August 5. Watch real-world containment rates, noisy-channel WER, and how Think Fast 2.0 stacks against GPT-Live outside vendor benches.