GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Arena Activates Battles-in-Direct Votes — Two-Month Backfill Could Shuffle Top-10 ELO

Arena’s “Battles in Direct” experiment launched quietly in March 2026: one in ten direct-chat sessions on the platform was converted into a two-model comparison, with users voting on preference. For two months, those votes were collected but kept out of the official leaderboard. That changed on May 12.

As of this week, Battles-in-Direct votes count toward ELO rankings. Over the next 30 days, Arena will backfill March through May data — meaning every model active since the experiment began is about to receive a large retroactive score update derived from a fundamentally different prompt distribution.

Why the Distribution Shift Matters

Standard Arena battles are short, discrete comparisons triggered by users visiting the platform specifically to run a head-to-head. Battles in Direct arrive mid-conversation — with prior context, longer queries, and task-specific framing. Arena’s May 12 changelog confirmed the distributional shift explicitly: “direct battles arrive with prior conversational context and skew toward longer queries, harder prompts, coding, and instruction following.”

That profile systematically benefits models with strong multi-turn capability and extended reasoning. The models most exposed to upside: Claude Opus 4.6 Thinking, GPT-5.5, Gemini 3.1 Pro Preview, and Kimi K2.6 — all of which have demonstrated outsized performance on precisely the tasks Battles in Direct over-represent. Models optimised for short, standalone responses may see their ELO diluted as the vote pool expands and harder prompts receive more weight.

One methodological guarantee matters here: sessions where users voted “skip” are excluded. Every vote in the backfill is a genuine stated preference, not a neutral signal.

The Current Top 10 Under Pressure

The Text arena top 10 at April close:

RankModelELO
1Claude Opus 4.6 Thinking1504
2Gemini 3.1 Pro Preview1493
3GPT-5.4 High1484
4Grok 4.201471
5DeepSeek V4 Pro1462
6Claude Sonnet 4.61458
7GPT-5.4 Standard1455
8Gemini 3.0 Pro1449
9Qwen 3.6-Plus1447
10Meta Muse Spark1441

Positions 2 through 10 span 63 ELO points. Most pairs in that range are within overlapping confidence intervals — which is to say, their relative ordering is not statistically resolved. The backfill will be the first significant re-ranking event since GPT-5.5-High entered in late April.

Confidence Intervals Will Tighten

The increased daily vote volume that Battles in Direct injects will tighten confidence intervals across all categories. Models currently tied within noise bands may see their positions resolve. That is particularly relevant in Code Arena and Vision Arena, where Arena’s July 2025 methodology update (frequency-based reweighting) widened confidence intervals by increasing variance, and where several pairs currently overlap.

Forward-Looking Impact

Models added to Arena after May 12 will have Battles-in-Direct counted from day one, giving newer entrants a structural volume advantage over older models whose pre-May direct battles were not captured. Any model launching on Arena in May or June will accumulate richer, harder-prompt vote data from the outset — which may accelerate the rate at which new entrants climb toward the top tier.

The next Arena leaderboard snapshot will be the first to incorporate this methodology change. Expect movement.