GPT-5.5-High Joins Arena at 1504 ELO — OpenAI Ties Anthropic at the Top for the First Time
GPT-5.5-high entered Chatbot Arena on April 27, 2026, posting 1504 on the main text leaderboard — the same score as Claude Opus 4.7. It is the first time an OpenAI model has reached the Arena summit since Anthropic’s Opus series took over in April.
On the code leaderboard, GPT-5.5-high hits 1558, ahead of Claude Opus 4.7’s code score. Arena also added GPT-5.5-high to Expert, Search, Document, and Vision leaderboards in the same push, giving it the broadest multi-domain footprint of any model added to Arena this month.
The Numbers
| Leaderboard | GPT-5.5-high ELO | Claude Opus 4.7 ELO |
|---|---|---|
| Text | 1504 | 1504 |
| Code | 1558 | ~1540 |
| Vision | 1312 | — |
The text tie is statistically meaningful. Arena’s confidence intervals at this score range are roughly ±10 ELO, so the overlap is real. Both models are operating at the frontier of what human preference voting can distinguish.
What GPT-5.5 Costs
GPT-5.5 is priced at $5 per million input tokens and $30 per million output tokens via the API — identical input pricing to Claude Opus 4.7 ($5/$25) but slightly higher on output. A Fast mode variant runs 1.5x faster at 2.5x the cost. Context window is 1 million tokens.
The proximity of the scores at equivalent price points makes the GPT-5.5 vs. Opus 4.7 decision a task-specific call rather than a headline performance one. GPT-5.5-high’s 1558 code score is the sharper differentiator for engineering teams choosing between the two.
Context
OpenAI’s own Terminal-Bench 2.0 results (82.0% with the Codex agent) established GPT-5.5’s agentic credentials three days before the Arena data arrived. Arena adds the human preference signal. The combined picture: GPT-5.5-high matches Opus 4.7 on open-ended conversation quality and leads on automated coding tasks when paired with OpenAI’s own scaffolding.
GPT-5.5’s hallucination rate remains a known liability. Artificial Analysis puts it at 86%, the highest of any frontier model — a real constraint for production workflows that require factual reliability over creative output.
The arena tie effectively resets the frontier comparison to cost structure, latency, and use-case fit. Neither lab has a clear head-to-head superiority at this price point.