GPT-5.5 xHigh Debuts at #2 on Agent Arena — 11.0% Composite, 3.2 Points Above GPT-5.5 High
GPT-5.5 xHigh entered Arena’s Agent leaderboard on June 11 with 7,376 votes logged — small compared to the 33,000-plus held by its High sibling, but directionally consistent. At 11.03% ± 2.52% composite, the xHigh variant sits at rank 2, between Claude Fable 5 and Claude Opus 4.8 Thinking. GPT-5.5 High in the same leaderboard sits at 7.80% (rank 5), a gap of 3.23 percentage points.
That gap matters. In real-world agent tasks, the difference between xHigh and High modes is not a marginal refinement — it shifts OpenAI’s best model from fifth to second place, leapfrogging two Anthropic models in the process.
Where It Lands
| Rank | Model | Composite | Votes |
|---|---|---|---|
| 1 | Claude Fable 5 (High) | 13.68% ± 1.59% | 16,280 |
| 2 | GPT-5.5 (xHigh) | 11.03% ± 2.52% | 7,376 |
| 3 | Claude Opus 4.8 (Thinking) | 9.05% ± 1.34% | 24,792 |
| 4 | Claude Opus 4.7 (Thinking) | 8.45% ± 1.26% | 28,379 |
| 5 | GPT-5.5 (High) | 7.80% ± 1.10% | 33,862 |
The composite scores here represent real user-evaluated task success across five categories — not synthetic benchmark scores. Even the top model completes fewer than 14% of tasks by this metric, which reflects the hardness of Agent Arena’s live task set rather than a ceiling on absolute capability.
The xHigh Effect
xHigh in GPT-5.5 designates maximum compute allocation per inference step — extended chain-of-thought, longer rollouts, no token budget constraint. For agentic trajectories, where planning over many steps compounds marginal improvements in reasoning quality, the effect is visible in the final numbers.
The 3.23-point gap between xHigh and High (11.03% vs 7.80%) is larger than the gap between High and Opus 4.7 Thinking (7.80% vs 8.45% — High actually trails Opus 4.7 Thinking at rank 4). Interpreted differently: turning xHigh off drops OpenAI’s best model below two Anthropic reasoning models.
The Gap to Fable 5
Claude Fable 5 (High) holds a 2.65-percentage-point lead over GPT-5.5 xHigh (13.68% vs 11.03%). Given xHigh’s early vote count (7,376 vs Fable 5’s 16,280), the margin has room to narrow or widen as the sample builds. What doesn’t change is Fable 5’s current rank-1 position — it has held the top spot since entering the leaderboard on June 10.
Sub-category breakdown for xHigh vs Fable 5 shows one area where they match exactly: 27.74% in one unspecified category. Fable 5 leads in four of five categories; xHigh leads in the category involving latency-tolerant longer tasks (14.67% vs 10.16% for Fable 5). That’s consistent with xHigh’s extended compute profile.
Context
Agent Arena uses live human evaluators to rate outputs across multi-step tasks. It is distinct from SWE-bench Verified or Terminal-Bench, both of which use deterministic grading on fixed task sets. Arena scores capture real-world preference and task completion as judged by humans, which makes them noisy but also harder to scaffold-optimise.
GPT-5.5 xHigh’s arrival at #2 confirms that OpenAI and Anthropic are now effectively trading positions at the top of the agentic leaderboard. Grok 4.3, which ranked last when the Agent Arena launched, remains at rank 22 with a 8.09% composite (High variant); the non-High Grok 4.3 sits at rank 25 last at 18.30% — that higher score reflecting a different balance of sub-category wins and losses in the Arena format.