GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Alibaba's Fun-Realtime-TTS Takes AA Speech Arena #1 at Elo 1,219 — Edges Google by 5 Points

Alibaba has the top voice AI model in the Artificial Analysis Speech Arena as of June 3. Fun-Realtime-TTS posted an Elo of 1,219 with 962 arena appearances, overtaking Google’s Gemini 3.1 Flash TTS at 1,214 and Inworld’s Realtime TTS-2 Research Preview at 1,209. Alibaba’s previous entry, Fun-Realtime-TTS-Preview, had reached #7. This is a 6-position jump and the first #1 for any Alibaba model on the speech leaderboard.

Leaderboard Snapshot

RankModelElo
1Fun-Realtime-TTS1,219
2Gemini 3.1 Flash TTS1,214
3Inworld Realtime TTS-2 Research Preview1,209
4Cartesia Sonic 3.51,203

The top five are separated by 24 Elo points — the tightest field in any Arena leaderboard this month. At this compression, a few hundred additional arena votes can reshuffle the rankings. The position is meaningful but not dominant.

Pricing and Capabilities

Fun-Realtime-TTS prices at $27.59/1M characters, placing it between Gemini 3.1 Flash TTS ($18.3/1M) and Inworld Realtime TTS 1.5 Max ($35/1M). It is more expensive than the Google option it displaced from #1 by a factor of 1.5x.

Feature coverage: real-time speech generation, voice cloning, voice design, multilingual output, and support for regional accents and dialects. The model is available via Alibaba Cloud API.

Context

The voice AI market is running a distinct competitive dynamic from the main LLM leaderboards. The top four positions are held by three labs — Alibaba, Google, and Inworld — rather than the OpenAI/Anthropic/Google axis that dominates text rankings. CartesIa holds a position with Sonic 3.5 at #4 despite being a specialist TTS company with no general LLM product.

The Speech Arena runs on human preference votes in blind matchups, which weights naturalness and prosody above the latency and multilingual accuracy metrics that enterprise TTS buyers typically care about. Arena position is a proxy for perceived quality rather than a comprehensive benchmark. That said, reaching #1 from #7 in one model generation is a real capability jump regardless of how you weight the methodology.

This is the second Chinese lab result this week to top a specialized Artificial Analysis leaderboard, following Alibaba’s HappyHorse-1.0 maintaining its lead in the Video Arena from earlier in the month. Alibaba’s model portfolio is increasingly covering the full stack of generative media outputs.