GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.5 xHigh Debuts at #2 on Agent Arena — 11.0% Composite, 3.2 Points Above GPT-5.5 High

GPT-5.5 xHigh entered Arena’s Agent leaderboard on June 11 with 7,376 votes logged — small compared to the 33,000-plus held by its High sibling, but directionally consistent. At 11.03% ± 2.52% composite, the xHigh variant sits at rank 2, between Claude Fable 5 and Claude Opus 4.8 Thinking. GPT-5.5 High in the same leaderboard sits at 7.80% (rank 5), a gap of 3.23 percentage points.

That gap matters. In real-world agent tasks, the difference between xHigh and High modes is not a marginal refinement — it shifts OpenAI’s best model from fifth to second place, leapfrogging two Anthropic models in the process.

Where It Lands

RankModelCompositeVotes
1Claude Fable 5 (High)13.68% ± 1.59%16,280
2GPT-5.5 (xHigh)11.03% ± 2.52%7,376
3Claude Opus 4.8 (Thinking)9.05% ± 1.34%24,792
4Claude Opus 4.7 (Thinking)8.45% ± 1.26%28,379
5GPT-5.5 (High)7.80% ± 1.10%33,862

The composite scores here represent real user-evaluated task success across five categories — not synthetic benchmark scores. Even the top model completes fewer than 14% of tasks by this metric, which reflects the hardness of Agent Arena’s live task set rather than a ceiling on absolute capability.

The xHigh Effect

xHigh in GPT-5.5 designates maximum compute allocation per inference step — extended chain-of-thought, longer rollouts, no token budget constraint. For agentic trajectories, where planning over many steps compounds marginal improvements in reasoning quality, the effect is visible in the final numbers.

The 3.23-point gap between xHigh and High (11.03% vs 7.80%) is larger than the gap between High and Opus 4.7 Thinking (7.80% vs 8.45% — High actually trails Opus 4.7 Thinking at rank 4). Interpreted differently: turning xHigh off drops OpenAI’s best model below two Anthropic reasoning models.

The Gap to Fable 5

Claude Fable 5 (High) holds a 2.65-percentage-point lead over GPT-5.5 xHigh (13.68% vs 11.03%). Given xHigh’s early vote count (7,376 vs Fable 5’s 16,280), the margin has room to narrow or widen as the sample builds. What doesn’t change is Fable 5’s current rank-1 position — it has held the top spot since entering the leaderboard on June 10.

Sub-category breakdown for xHigh vs Fable 5 shows one area where they match exactly: 27.74% in one unspecified category. Fable 5 leads in four of five categories; xHigh leads in the category involving latency-tolerant longer tasks (14.67% vs 10.16% for Fable 5). That’s consistent with xHigh’s extended compute profile.

Context

Agent Arena uses live human evaluators to rate outputs across multi-step tasks. It is distinct from SWE-bench Verified or Terminal-Bench, both of which use deterministic grading on fixed task sets. Arena scores capture real-world preference and task completion as judged by humans, which makes them noisy but also harder to scaffold-optimise.

GPT-5.5 xHigh’s arrival at #2 confirms that OpenAI and Anthropic are now effectively trading positions at the top of the agentic leaderboard. Grok 4.3, which ranked last when the Agent Arena launched, remains at rank 22 with a 8.09% composite (High variant); the non-High Grok 4.3 sits at rank 25 last at 18.30% — that higher score reflecting a different balance of sub-category wins and losses in the Arena format.