Arena Adds Battles-in-Direct to ELO — Position Bias and Org-Context Advantage Now Corrected
Arena updated its leaderboard methodology on May 12, making a change that will affect how every frontier model is ranked over the next 30 days.
Since March 2026, 10% of direct chat sessions on Arena have been silently converted into head-to-head battles between two randomly sampled models. Those votes now feed into leaderboards — starting with new votes today, with a backfill of March-May data rolling out over the next month.
What Changes
The direct-battle corpus is different from Arena’s standard blind-battle mode. Prompts arrive with prior conversational context and skew toward longer queries, harder tasks, and coding and instruction-following challenges. Standard Arena battles are typically cold-start, single-turn interactions.
The practical effect: daily vote volume rises, confidence intervals tighten, and rankings stabilise faster. Every direct-battle vote is a decisive A-or-B preference — Arena excludes sessions where the user votes skip.
Two Biases Discovered and Corrected
Analysis of the direct-battle data surfaced two systematic biases not present in standard mode:
Position bias — Model A (left-side model) received more votes independent of response quality. Likely driven by default click behaviour under cognitive load.
Org-context advantage — Models sharing an organisation with the prior conversational context received a scoring bump unrelated to actual response quality. A user deep in an Anthropic-context session would subtly favour Claude variants; an OpenAI-context session would subtly favour GPT variants.
Both corrections are applied in the Bradley-Terry fitting using two new binary features — is_direct_battle and same_org_indicator — whose coefficients are learned during fitting to absorb each bias. The corrections apply to text, vision, search, and document arenas.
Implications for the Ticker
Models optimised for extended, multi-turn dialogue should benefit as direct-battle votes are backfilled. Models tuned primarily for cold-start single-turn tasks may see marginal ELO pressure.
The org-context correction specifically affects models from labs that dominate a user’s session history. Anthropic and OpenAI models — which have the highest session retention — are most exposed to this correction. Whether that nets as positive or negative depends on whether their models were benefiting from the org-context bias or losing to it.
Expect non-trivial rank movement through the first two weeks of June as March-May backfill data flows in. The confidence interval narrowing will make those movements stick faster than usual.
Recent Additions (May 6-8, 2026)
Separate from the methodology change, Arena added several models this week:
gpt-5.5-instantjoined Text, Vision, and Document leaderboards (May 6 and May 8)ernie-5.1added to Search leaderboard (May 8)gemma-4-31bandgemma-4-26b-a4badded to Vision leaderboard (May 7)qwen3.6-max-previewandhunyuan-hy3-previewadded to Code leaderboard (May 7)grok-imagine-image-qualityadded to Text-to-Image and Image Edit leaderboards (May 6)uni-1.1-maxanduni-1.1added to Text-to-Image and Image Edit leaderboards (May 5)