GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Arena Adds Battles-in-Direct to ELO — Position Bias and Org-Context Advantage Now Corrected

Arena updated its leaderboard methodology on May 12, making a change that will affect how every frontier model is ranked over the next 30 days.

Since March 2026, 10% of direct chat sessions on Arena have been silently converted into head-to-head battles between two randomly sampled models. Those votes now feed into leaderboards — starting with new votes today, with a backfill of March-May data rolling out over the next month.

What Changes

The direct-battle corpus is different from Arena’s standard blind-battle mode. Prompts arrive with prior conversational context and skew toward longer queries, harder tasks, and coding and instruction-following challenges. Standard Arena battles are typically cold-start, single-turn interactions.

The practical effect: daily vote volume rises, confidence intervals tighten, and rankings stabilise faster. Every direct-battle vote is a decisive A-or-B preference — Arena excludes sessions where the user votes skip.

Two Biases Discovered and Corrected

Analysis of the direct-battle data surfaced two systematic biases not present in standard mode:

Position bias — Model A (left-side model) received more votes independent of response quality. Likely driven by default click behaviour under cognitive load.

Org-context advantage — Models sharing an organisation with the prior conversational context received a scoring bump unrelated to actual response quality. A user deep in an Anthropic-context session would subtly favour Claude variants; an OpenAI-context session would subtly favour GPT variants.

Both corrections are applied in the Bradley-Terry fitting using two new binary features — is_direct_battle and same_org_indicator — whose coefficients are learned during fitting to absorb each bias. The corrections apply to text, vision, search, and document arenas.

Implications for the Ticker

Models optimised for extended, multi-turn dialogue should benefit as direct-battle votes are backfilled. Models tuned primarily for cold-start single-turn tasks may see marginal ELO pressure.

The org-context correction specifically affects models from labs that dominate a user’s session history. Anthropic and OpenAI models — which have the highest session retention — are most exposed to this correction. Whether that nets as positive or negative depends on whether their models were benefiting from the org-context bias or losing to it.

Expect non-trivial rank movement through the first two weeks of June as March-May backfill data flows in. The confidence interval narrowing will make those movements stick faster than usual.

Recent Additions (May 6-8, 2026)

Separate from the methodology change, Arena added several models this week:

  • gpt-5.5-instant joined Text, Vision, and Document leaderboards (May 6 and May 8)
  • ernie-5.1 added to Search leaderboard (May 8)
  • gemma-4-31b and gemma-4-26b-a4b added to Vision leaderboard (May 7)
  • qwen3.6-max-preview and hunyuan-hy3-preview added to Code leaderboard (May 7)
  • grok-imagine-image-quality added to Text-to-Image and Image Edit leaderboards (May 6)
  • uni-1.1-max and uni-1.1 added to Text-to-Image and Image Edit leaderboards (May 5)