GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

GPT-5.5-High Enters All Six Arena Categories — OpenAI and Anthropic Now Head-to-Head Everywhere

Chatbot Arena runs six competitive leaderboards: Text, Code, Expert, Search, Document, and Vision. On April 27, GPT-5.5-high was added to all six simultaneously — the first time a single model from any lab has achieved full-category coverage in a single day. Claude Opus 4.7 was added to the Search leaderboard on the same date.

What Changed

GPT-5.5-high had previously been added to the text Arena shortly after launch, posting an ELO of 1504. The April 27 expansion moved it into every modality track Arena runs — Code, Expert, Search, Document, and Vision in addition to Text.

Claude Opus 4.7, already live in Text, Code, and Expert, now also has a Search leaderboard entry. Search tracks real-time information retrieval capability, testing how models handle web grounding and temporal recency — one of the sharper functional differentiators between frontier labs.

Also on April 23: DeepSeek V4 Pro, DeepSeek V4 Pro Thinking, and DeepSeek V4 Flash Thinking joined the Text and Code leaderboards. Qwen Image 2.0 Pro entered Text-to-Image and Image Edit. Arena is currently absorbing four major model families across modality tracks simultaneously.

The Intelligence Index Context

Artificial Analysis’s current rankings show GPT-5.5 (xhigh) leading at an Intelligence Index of 60, GPT-5.5 (high) at 59, followed by a three-way tie at 57 between Claude Opus 4.7 (max), Gemini 3.1 Pro Preview, and GPT-5.4 (xhigh). Open weights models trail: Kimi K2.6 at 54, DeepSeek V4 Pro at 52, GLM-5.1 at 51.

GPT-5.5-high, at index 59, is the second-best GPT-5.5 variant. OpenAI’s choice to push high rather than xhigh across all Arena categories first suggests coverage speed is the priority — not peak benchmark positioning.

Why Full-Spectrum Arena Coverage Matters

Arena ELO aggregates real user preference judgements, not controlled lab evaluations. Getting listed across all six categories creates comparative data in Search, Document, and Vision where GPT-5.5 had no public performance record before. For labs and enterprise buyers, Arena’s multi-category coverage is increasingly the standard reference for model selection.

For Anthropic, the Search addition for Opus 4.7 is worth watching. Search Arena rewards web retrieval accuracy and freshness — a category where Anthropic’s flagship models have historically tracked below OpenAI variants. Initial ELO for Opus 4.7 in Search is not yet published; ratings will emerge as battle volume accrues.

Expect meaningful ELO data across most GPT-5.5-high categories within two to three weeks. The Code and Expert tracks typically accumulate battles faster than Document and Vision, so those two will produce stable rankings first.

The full coverage picture creates, for the first time, a head-to-head record between GPT-5.5-high and Claude Opus 4.7 across every domain Arena can measure.