Grok Voice Think Fast 2.0 Leads AA Speech Arena at 82.9% — Beats GPT-Realtime by 3.8 Points
xAI released Grok Voice Think Fast 2.0 on July 29, the successor to the voice model that shipped in May and posted 70% autonomous resolution at Starlink internal deployments. The new model leads the Artificial Analysis Speech-to-Speech Quality Index at 82.9%, ahead of GPT-Realtime-2.1 High at 79.1% and Gemini 3.1 Flash High at 69.5%.
Benchmark Numbers
The four-model comparison xAI published against Artificial Analysis data:
| Benchmark | Grok VTF 2.0 | Grok VTF 1.0 | GPT-Realtime-2.1 High | Gemini 3.1 Flash High |
|---|---|---|---|---|
| AA S2S Quality Index | 82.9% | 75.7% | 79.1% | 69.5% |
| Big Bench Audio (Speech Reasoning) | 97.2% | 97.1% | 96.0% | 96.6% |
| Full Duplex Bench (Conversational Dynamics) | 95.1% | 77.8% | 95.7% | 74.3% |
| τ-voice Bench (Agentic Performance) | 56.5% | 52.1% | 45.7% | 37.7% |
| Time to First Audio | 0.70s | 1.25s | — | 2.98s |
The overall index advantage comes primarily from τ-voice Bench agentic performance (56.5% vs 45.7% for GPT-Realtime-2.1) and the latency gap. Full Duplex Bench is the one dimension where GPT-Realtime-2.1 High edges ahead by 0.6 points.
Speech Reasoning and Transcription
Speech reasoning scores are compressed at the top: all four models score within 1.2 percentage points on Big Bench Audio. The competitive separation is happening on agentic capability — τ-voice Bench’s measure of autonomous task completion — and on raw latency.
xAI separately reports a 1.5-2.0x improvement in transcription accuracy across 24 languages compared to Grok VTF 1.0, evaluated on thousands of short phrases. The transcription claim covers the same language set as Grok Voice’s API release in June, which included speech-to-text at $0.10/hour.
Why τ-voice Bench Matters
τ-voice Bench is the agentic performance dimension of the AA Speech leaderboard, measuring how well voice models complete autonomous tool-use tasks in conversation — not just how accurately they transcribe or how naturally they respond. At 56.5%, Grok VTF 2.0 posts the highest score on this dimension of any model currently on the AA Speech-to-Speech leaderboard. GPT-Realtime-2.1 High sits 10.8 points back at 45.7%.
This is the voice benchmark most relevant to the enterprise call center and autonomous agent market that xAI has been targeting since the Starlink deployment. Transcription accuracy and conversational flow matter for general use; agentic performance is what determines whether a voice model can close tickets, complete transactions, or route queries without human handoff.
Context
Grok VTF 1.0 shipped in May with enterprise deployment data from Starlink showing 70% autonomous resolution. The 2.0 model does not include an updated autonomous resolution figure from a production deployment, which is the metric that will matter most for enterprise evaluation. The benchmark numbers establish where the model sits in the competitive landscape; real-world task completion rates at scale will be the test that determines adoption.
The 0.70-second time to first audio is the fastest of any model in the comparison. For live conversational applications — customer service, voice agents, real-time translation — latency below one second is a practical threshold that meaningfully changes user experience. Grok VTF 2.0 is the only model in the AA Speech leaderboard currently below it.