GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Artificial Analysis Ships Speech-to-Speech Index: GPT-Realtime-2 Leads at 77.2%, Grok Voice Tops Two of Three Dimensions

Artificial Analysis published a Speech-to-Speech Index on June 23 that combines three existing benchmarks into a single composite score for native audio models. The index covers Speech Reasoning, Conversational Dynamics, and Agentic Performance, weighted equally. Models must post valid results across all three to appear in the ranking.

Four models are currently included. GPT-Realtime-2 (High) leads overall. Grok Voice Think Fast 1.0 leads two of the three dimensions.

Index Scores

ModelOverallCost/hr
GPT-Realtime-2 (High)77.2%$4.14
Grok Voice Think Fast 1.075.7%$3.00
GPT-Realtime-1.572.0%-
Gemini 3.1 Flash Live Preview (High)69.5%$1.75
Deepslate Opal62.1%-

The 1.5-point gap between GPT-Realtime-2 and Grok Voice Think Fast is narrow on the composite, but the dimension breakdown tells a different story.

What Each Dimension Shows

Speech Reasoning (Big Bench Audio): 1,000 questions across Formal Fallacies, Navigate, Object Counting, and Web of Lies. Scores are tightly clustered at the top. Grok Voice Think Fast 1.0 leads at 97.1%.

Conversational Dynamics (Full Duplex Bench): Measures pause handling, turn-taking, interruption, and backchannel responses. GPT-Realtime-2 leads overall, but the specific standout is GPT-Realtime-2 (Minimal), which tops this dimension at 96.1%.

Agentic Performance (τ-Voice): End-to-end customer service task completion across Airline, Retail, and Telecom. This is the hardest dimension by a significant margin. Grok Voice Think Fast 1.0 leads at 52.1%, followed by GPT-Realtime-2 (High) at 39.8%. No model currently scores above 53%. The 12.3-point gap between first and second on this dimension is the widest spread in the index.

Agentic Performance is the differentiator that matters most for production voice deployments. A model that handles backchannel cues well but fails task completion is a conversational party trick, not a deployable agent. On that measure, Grok Voice Think Fast 1.0 is meaningfully ahead.

Speed and Cost

Deepslate Opal has the fastest time to first audio at 0.44 seconds, well ahead of GPT-Realtime-1.5 at 0.82s and Grok Voice Think Fast 1.0 at 1.25s. GPT-Realtime-2 (High) takes 2.33 seconds; Gemini 3.1 Flash Live Preview (High) takes 2.98 seconds.

Speed and quality are inversely correlated across the current field. The fastest model (Deepslate Opal at 0.44s TTFA) scores 62.1% on the composite. The highest-scoring model (GPT-Realtime-2 at 77.2%) takes 5x longer to start speaking. For interactive voice applications, that latency tradeoff may override the quality ranking.

Cost-per-hour: Gemini 3.1 Flash Live Preview (Minimal) is cheapest at $1.50, scoring 56.6%. Gemini 3.1 Flash Live Preview (High) is $1.75 at 69.5%, making it the strongest cost-adjusted option currently in the index. Grok Voice Think Fast 1.0 at $3.00 and GPT-Realtime-2 (High) at $4.14 compete on raw capability, not price.

The Larger Picture

Native speech-to-speech models have moved faster than the benchmark coverage supporting them. The AA index brings three existing evaluation frameworks into a single ranking for the first time. Agentic Performance being the hardest dimension, and the only one where no model cracks 53%, indicates the gap between conversational capability and task completion remains large for voice-first AI.

AA says it will continue adding models to the index as results are submitted.