Arena Activates Battles-in-Direct Votes — Two-Month Backfill Could Shuffle Top-10 ELO
Arena’s “Battles in Direct” experiment launched quietly in March 2026: one in ten direct-chat sessions on the platform was converted into a two-model comparison, with users voting on preference. For two months, those votes were collected but kept out of the official leaderboard. That changed on May 12.
As of this week, Battles-in-Direct votes count toward ELO rankings. Over the next 30 days, Arena will backfill March through May data — meaning every model active since the experiment began is about to receive a large retroactive score update derived from a fundamentally different prompt distribution.
Why the Distribution Shift Matters
Standard Arena battles are short, discrete comparisons triggered by users visiting the platform specifically to run a head-to-head. Battles in Direct arrive mid-conversation — with prior context, longer queries, and task-specific framing. Arena’s May 12 changelog confirmed the distributional shift explicitly: “direct battles arrive with prior conversational context and skew toward longer queries, harder prompts, coding, and instruction following.”
That profile systematically benefits models with strong multi-turn capability and extended reasoning. The models most exposed to upside: Claude Opus 4.6 Thinking, GPT-5.5, Gemini 3.1 Pro Preview, and Kimi K2.6 — all of which have demonstrated outsized performance on precisely the tasks Battles in Direct over-represent. Models optimised for short, standalone responses may see their ELO diluted as the vote pool expands and harder prompts receive more weight.
One methodological guarantee matters here: sessions where users voted “skip” are excluded. Every vote in the backfill is a genuine stated preference, not a neutral signal.
The Current Top 10 Under Pressure
The Text arena top 10 at April close:
| Rank | Model | ELO |
|---|---|---|
| 1 | Claude Opus 4.6 Thinking | 1504 |
| 2 | Gemini 3.1 Pro Preview | 1493 |
| 3 | GPT-5.4 High | 1484 |
| 4 | Grok 4.20 | 1471 |
| 5 | DeepSeek V4 Pro | 1462 |
| 6 | Claude Sonnet 4.6 | 1458 |
| 7 | GPT-5.4 Standard | 1455 |
| 8 | Gemini 3.0 Pro | 1449 |
| 9 | Qwen 3.6-Plus | 1447 |
| 10 | Meta Muse Spark | 1441 |
Positions 2 through 10 span 63 ELO points. Most pairs in that range are within overlapping confidence intervals — which is to say, their relative ordering is not statistically resolved. The backfill will be the first significant re-ranking event since GPT-5.5-High entered in late April.
Confidence Intervals Will Tighten
The increased daily vote volume that Battles in Direct injects will tighten confidence intervals across all categories. Models currently tied within noise bands may see their positions resolve. That is particularly relevant in Code Arena and Vision Arena, where Arena’s July 2025 methodology update (frequency-based reweighting) widened confidence intervals by increasing variance, and where several pairs currently overlap.
Forward-Looking Impact
Models added to Arena after May 12 will have Battles-in-Direct counted from day one, giving newer entrants a structural volume advantage over older models whose pre-May direct battles were not captured. Any model launching on Arena in May or June will accumulate richer, harder-prompt vote data from the outset — which may accelerate the rate at which new entrants climb toward the top tier.
The next Arena leaderboard snapshot will be the first to incorporate this methodology change. Expect movement.