GPT-5.5-High Enters All Six Arena Categories — OpenAI and Anthropic Now Head-to-Head Everywhere
Chatbot Arena runs six competitive leaderboards: Text, Code, Expert, Search, Document, and Vision. On April 27, GPT-5.5-high was added to all six simultaneously — the first time a single model from any lab has achieved full-category coverage in a single day. Claude Opus 4.7 was added to the Search leaderboard on the same date.
What Changed
GPT-5.5-high had previously been added to the text Arena shortly after launch, posting an ELO of 1504. The April 27 expansion moved it into every modality track Arena runs — Code, Expert, Search, Document, and Vision in addition to Text.
Claude Opus 4.7, already live in Text, Code, and Expert, now also has a Search leaderboard entry. Search tracks real-time information retrieval capability, testing how models handle web grounding and temporal recency — one of the sharper functional differentiators between frontier labs.
Also on April 23: DeepSeek V4 Pro, DeepSeek V4 Pro Thinking, and DeepSeek V4 Flash Thinking joined the Text and Code leaderboards. Qwen Image 2.0 Pro entered Text-to-Image and Image Edit. Arena is currently absorbing four major model families across modality tracks simultaneously.
The Intelligence Index Context
Artificial Analysis’s current rankings show GPT-5.5 (xhigh) leading at an Intelligence Index of 60, GPT-5.5 (high) at 59, followed by a three-way tie at 57 between Claude Opus 4.7 (max), Gemini 3.1 Pro Preview, and GPT-5.4 (xhigh). Open weights models trail: Kimi K2.6 at 54, DeepSeek V4 Pro at 52, GLM-5.1 at 51.
GPT-5.5-high, at index 59, is the second-best GPT-5.5 variant. OpenAI’s choice to push high rather than xhigh across all Arena categories first suggests coverage speed is the priority — not peak benchmark positioning.
Why Full-Spectrum Arena Coverage Matters
Arena ELO aggregates real user preference judgements, not controlled lab evaluations. Getting listed across all six categories creates comparative data in Search, Document, and Vision where GPT-5.5 had no public performance record before. For labs and enterprise buyers, Arena’s multi-category coverage is increasingly the standard reference for model selection.
For Anthropic, the Search addition for Opus 4.7 is worth watching. Search Arena rewards web retrieval accuracy and freshness — a category where Anthropic’s flagship models have historically tracked below OpenAI variants. Initial ELO for Opus 4.7 in Search is not yet published; ratings will emerge as battle volume accrues.
Expect meaningful ELO data across most GPT-5.5-high categories within two to three weeks. The Code and Expert tracks typically accumulate battles faster than Document and Vision, so those two will produce stable rankings first.
The full coverage picture creates, for the first time, a head-to-head record between GPT-5.5-high and Claude Opus 4.7 across every domain Arena can measure.