GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Arena Adds Six Models Across Four Leaderboards in 48 Hours — Broadest Coverage Expansion of 2026

Arena added six models across four leaderboard categories between May 7 and May 8 — the single largest two-day expansion to the platform’s coverage in 2026.

What Entered and Where

May 8:

  • GPT-5.5-Instant → Text, Vision, and Document leaderboards
  • ERNIE 5.1 → Search leaderboard

May 7:

  • Gemma 4-31B → Vision leaderboard
  • Gemma 4-26B-A4B → Vision leaderboard
  • Qwen 3.6 Max Preview → Code leaderboard
  • HunYuan Hy3-Preview → Code leaderboard

Why GPT-5.5-Instant’s Entry Is the One to Watch

GPT-5.5-Instant became the default ChatGPT model in early May. It carries a lower price point than GPT-5.5-High — $5/M input vs the High tier — and was positioned by OpenAI as the everyday response model with meaningfully fewer hallucinations than earlier o-series variants.

Its Arena entry is the first time the model sits in the same blind preference evaluation as Claude Opus 4.7, Gemini 3.1 Pro Preview, and GPT-5.5-High. Text, Vision, and Document coverage across all three categories means it will accumulate votes across the full range of use cases OpenAI intends for it. ELO scores for Instant against the current frontier — Opus 4.7 and GPT-5.5-High are tied around 1505 — will resolve the question of how much capability OpenAI’s cost optimisation removed.

Baidu’s ERNIE 5.1 entering the Search leaderboard matters for a different reason. Earlier analysis placed ERNIE 5.1 at Arena #13 globally on the general text leaderboard, with a lead in Legal tasks worldwide. The Search category is a separate evaluation surface where native grounding and citation accuracy are what voters judge. ERNIE’s position in legal contexts suggests strong structured retrieval, which should translate well here — but the crowd preference data will either confirm or cut that assumption.

The Code Category Gets Competitive

HunYuan Hy3-Preview brings 295B total parameters with only 21B active — a 74.4% SWE-bench Verified result positions it at the lower-mid frontier tier. Qwen 3.6 Max Preview enters having scored 52 on Artificial Analysis’s Intelligence Index, placing it at the top of every Chinese model on that scale.

Both enter a Code leaderboard currently dominated by GPT-5.5-High and Opus 4.7. The Code category is where preference gaps between models are most actionable for developers. Their head-to-head ELO against established entrants in the next few weeks will be the practical signal.

The Gemma 4 Vision Data

Gemma 4-31B and the 26B-A4B mixture-of-experts variant enter Vision without the same buzz as the frontier models but with a specific audience: teams evaluating open-weights options for vision tasks. The 26B-A4B runs on a fraction of the memory footprint of the 31B dense model. Arena’s Vision ELO will give the first human preference signal on whether that parameter efficiency comes at a noticeable quality cost.

Leaderboard State

The core frontier standings have not changed with these additions. GPT-5.5-High and Claude Opus 4.7 are tied around 1505 ELO at the top; Gemini 3.1 Pro Preview sits just behind. The six new entrants are competing for mid-tier positions. What the expansion does is fill in the comparison matrix — making it possible to evaluate cost/performance tradeoffs for the full set without relying solely on provider-reported benchmarks.