Sunday, Aug 2 | --:--
Back to home

SpaceXAI Ships Grok Voice Think Fast 2.0 With Top Speech-to-Speech Scores

On July 29, 2026, SpaceXAI released Grok Voice Think Fast 2.0—its next-generation speech-to-speech voice model—claiming 82.9% on Artificial Analysis’s Speech-to-Speech Quality Index, 0.70s time-to-first-audio, $0.08/min pricing, and a default swap for grok-voice-latest on August 5.

Tech Insights Reporter 4 min read San Francisco, CA
Cover illustration for SpaceXAI Ships Grok Voice Think Fast 2.0 With Top Speech-to-Speech Scores

TLDR

SpaceXAI on July 29, 2026 announced Grok Voice Think Fast 2.0, its most capable speech-to-speech voice model. The company primary cites Artificial Analysis leadership on speech-to-speech quality (82.9% vs 79.1% for GPT-Realtime-2.1 High), 0.70s time to first audio (down from 1.25s on 1.0), roughly 0.4× the reasoning tokens of its predecessor in production, and $0.08 per minute of audio. grok-voice-latest migrates to 2.0 on August 5, 2026 unless pinned.

What shipped

Item Detail (x.ai primary)
Model Grok Voice Think Fast 2.0
Focus Speech reasoning, transcription accuracy, conversational tool use
AA Speech-to-Speech Quality 82.9% (1.0: 75.7%; GPT-Realtime-2.1 High: 79.1%)
tau-Voice agentic 56.5% vs 45.7% for GPT-Realtime-2.1 High (lab table)
Latency 0.70s time to first audio (1.0: 1.25s)
Pricing $0.08 / min of audio
Default cutover grok-voice-latest → 2.0 on Aug 5, 2026 (pin grok-voice-think-fast-1.0 to stay)
Production claim A/B on Starlink support/sales lines: higher conversion and containment

SpaceXAI positions 2.0 as reasoning while speaking—tool calls often fire before the first sentence finishes—without adding turn-taking latency. Transcription claims include 1.5–2.0× lower word-error relative to Deepgram Nova 3 and ElevenLabs Scribe v2 on internal multilingual short-phrase tests, with larger gaps under noise/telephony compression.

Independent rankings

Source Metric Snapshot
Artificial Analysis Speech-to-Speech Quality Index 82.9% Cited on x.ai launch post as leading GPT-Realtime-2.1 High (79.1%) and Gemini 3.1 Flash High (69.5%)
AA / lab tau-Voice Agentic voice 56.5% Lab table vs 45.7% (GPT-Realtime-2.1 High)

Product-line placement

Distinct from Grok 4.5 (text/coding flagship), Build Mode, and Imagine media models. Continues the Voice Agent API / Think Fast line launched earlier in 2026 (1.0). Competes with OpenAI GPT-Live / Realtime and Google Gemini voice stacks on full-duplex speech agents.

Why this story matters

Voice is becoming a primary enterprise surface for support, sales, and agent workflows—not a demo. Shipping a measurable step-change on third-party speech-to-speech boards and a default API migration date forces teams to re-evaluate cost, latency, and pin strategy before August 5. Watch real-world containment rates, noisy-channel WER, and how Think Fast 2.0 stacks against GPT-Live outside vendor benches.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading