GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Grok Voice Think Fast 2.0 Leads AA Speech Arena at 82.9% — Beats GPT-Realtime by 3.8 Points

xAI released Grok Voice Think Fast 2.0 on July 29, the successor to the voice model that shipped in May and posted 70% autonomous resolution at Starlink internal deployments. The new model leads the Artificial Analysis Speech-to-Speech Quality Index at 82.9%, ahead of GPT-Realtime-2.1 High at 79.1% and Gemini 3.1 Flash High at 69.5%.

Benchmark Numbers

The four-model comparison xAI published against Artificial Analysis data:

BenchmarkGrok VTF 2.0Grok VTF 1.0GPT-Realtime-2.1 HighGemini 3.1 Flash High
AA S2S Quality Index82.9%75.7%79.1%69.5%
Big Bench Audio (Speech Reasoning)97.2%97.1%96.0%96.6%
Full Duplex Bench (Conversational Dynamics)95.1%77.8%95.7%74.3%
τ-voice Bench (Agentic Performance)56.5%52.1%45.7%37.7%
Time to First Audio0.70s1.25s2.98s

The overall index advantage comes primarily from τ-voice Bench agentic performance (56.5% vs 45.7% for GPT-Realtime-2.1) and the latency gap. Full Duplex Bench is the one dimension where GPT-Realtime-2.1 High edges ahead by 0.6 points.

Speech Reasoning and Transcription

Speech reasoning scores are compressed at the top: all four models score within 1.2 percentage points on Big Bench Audio. The competitive separation is happening on agentic capability — τ-voice Bench’s measure of autonomous task completion — and on raw latency.

xAI separately reports a 1.5-2.0x improvement in transcription accuracy across 24 languages compared to Grok VTF 1.0, evaluated on thousands of short phrases. The transcription claim covers the same language set as Grok Voice’s API release in June, which included speech-to-text at $0.10/hour.

Why τ-voice Bench Matters

τ-voice Bench is the agentic performance dimension of the AA Speech leaderboard, measuring how well voice models complete autonomous tool-use tasks in conversation — not just how accurately they transcribe or how naturally they respond. At 56.5%, Grok VTF 2.0 posts the highest score on this dimension of any model currently on the AA Speech-to-Speech leaderboard. GPT-Realtime-2.1 High sits 10.8 points back at 45.7%.

This is the voice benchmark most relevant to the enterprise call center and autonomous agent market that xAI has been targeting since the Starlink deployment. Transcription accuracy and conversational flow matter for general use; agentic performance is what determines whether a voice model can close tickets, complete transactions, or route queries without human handoff.

Context

Grok VTF 1.0 shipped in May with enterprise deployment data from Starlink showing 70% autonomous resolution. The 2.0 model does not include an updated autonomous resolution figure from a production deployment, which is the metric that will matter most for enterprise evaluation. The benchmark numbers establish where the model sits in the competitive landscape; real-world task completion rates at scale will be the test that determines adoption.

The 0.70-second time to first audio is the fastest of any model in the comparison. For live conversational applications — customer service, voice agents, real-time translation — latency below one second is a practical threshold that meaningfully changes user experience. Grok VTF 2.0 is the only model in the AA Speech leaderboard currently below it.