GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Grok Voice Transcribe 2.0 Takes #1 on Artificial Analysis Streaming Leaderboard Among 32 Models

SpaceXAI released Grok Voice Transcribe 2.0 on September 18, claiming the top accuracy position on the Artificial Analysis streaming transcription leaderboard across all 32 streaming systems tracked on the platform, overtaking Meta’s Muse Voice Transcribe which had held the top spot.

The headline stat: word error rate cut in half versus Grok Voice Transcribe 1.0, with no price increase. Batch transcription stays at $0.10 per hour of audio; streaming stays at $0.20 per hour. Diarization, timestamps, and key terms are included at both tiers.

Where it matters

SpaceXAI’s pitch is that the model was built specifically for the hardest audio — telephone lines, overlapping speakers, regional accents, and credentials read aloud (account codes, email addresses). Internal evaluations run against four production traffic datasets: telephony, Grok conversations, spoken credentials, and multilingual voice commands. Grok Voice Transcribe 2.0 improved on 1.0 across all four, with the largest relative gap on multilingual voice commands (WER dropped from 20.6% to 6.8%), while telephony audio fell from 10.6% to 7.1%.

The model runs on the same audio foundation as Grok Voice, which already handles tens of thousands of customer-support calls daily, transcribes millions of hours of video narration, and powers the Grok assistant in Tesla vehicles.

Atlassian Loom is named as a new live customer for Grok Voice Transcribe 2.0 at launch.

Transition path

Grok Voice Transcribe 1.0 will be deprecated “in the coming weeks.” Developers who need to stay on 1.0 through the transition can pin grok-voice-transcribe-1.0 explicitly. Once Grok Voice Transcribe 2.0 becomes the Speech-to-Text API default, pinning will be the only path back.

Competitive context

The AA leaderboard places Grok Voice Transcribe 2.0 above Muse Voice Transcribe and Scribe v2 R on streaming accuracy. The accuracy-vs-price chart puts it in the top-right quadrant — best accuracy among low-cost options — at a price point well below several competitors with higher WER.

For enterprise buyers choosing a transcription vendor, the combination of first-place streaming accuracy, flat pricing, and an existing fleet that already handles real-world audio at scale is a credible pitch. Whether independent evaluation confirms the AA accuracy ranking will be the next test.