GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Claude Fable 5 Takes Agent Arena #1 With 12.94% Composite — GPT-5.5 High Falls to Fifth

Claude Fable 5 (High) entered Arena’s Agent Arena leaderboard on June 10 and immediately overtook GPT-5.5 (High), which had held first place for over two months. As of Friday, Fable 5’s composite score stands at 12.94% across 14,177 live user sessions — a 59% margin over GPT-5.5 (High)‘s current 8.16% and larger than any gap between first and the rest of the field since the leaderboard launched.

Current Standings

RankModelCompositeVotes
1Claude Fable 5 (High) ↑ NEW12.94% ± 2.10%14,177
2GPT-5.5 (xHigh) ↑ NEW10.56% ± 3.05%5,116
3Claude Opus 4.8 (Thinking)9.26% ± 1.41%22,491
4Claude Opus 4.7 (Thinking)8.64% ± 1.21%27,821
5GPT-5.5 (High)8.16% ± 1.12%31,694
6Claude Opus 4.67.98% ± 1.23%27,977
7Claude Opus 4.77.60% ± 1.26%27,958
8GPT-5.4 (High)7.29% ± 1.18%31,741
9GPT-5.57.10% ± 1.09%32,040
10Claude Opus 4.84.77% ± 1.76%20,031
11Claude Sonnet 4.63.39% ± 1.14%27,862

Agent Arena measures five real-world signals across live user sessions in Arena’s Agent Mode: task success, steerability, error recovery, user satisfaction, and tool fidelity. The composite uses causal inference methodology to isolate each model’s contribution; a higher score is a larger win margin in pairwise head-to-head tasks.

What Drives the Gap

Fable 5’s lead is not evenly distributed across the five sub-dimensions. Error recovery is the clearest separation: Fable 5 records 28.86% (±7.97%) versus Claude Opus 4.8 (Thinking)‘s 17.86% and GPT-5.5 (High)‘s 11.07%. Task success (first sub-dimension) follows: 17.55% for Fable 5 against 9.88% for Opus 4.8 (Thinking) and 6.40% for GPT-5.5 (High).

Error recovery is the more structurally significant signal. Agent tasks involve tool chains and mid-session corrections; a model that recovers from failed steps without user intervention completes more jobs and generates fewer escalations. That category is where Fable 5’s separation from the field is widest by absolute margin.

GPT-5.5 (High)‘s Directional Drop

GPT-5.5 (High) recorded 9.11% when the June 10 data was sampled. It now sits at 8.16% with 31,694 votes — a 0.95 percentage point drop on a high-confidence vote base. Its confidence interval of ±1.12% means the 9.11% score it held two days ago now sits outside that range. This is not a sampling artefact; the score moved down with more data, not up.

GPT-5.5 (xHigh) entered at rank 2 with 10.56%, but 5,116 votes gives it a ±3.05% band. That score is not yet directly comparable to Fable 5’s on statistical grounds.

Anthropic’s Structural Hold

Anthropic now holds seven of the top eleven Agent Arena positions: 1, 3, 4, 6, 7, 10, and 11. OpenAI places second and fifth through ninth. The thinking-mode divide that appeared when Opus 4.8 was added persists and has widened: the top four positions are all thinking-enabled models, and the only non-thinking models in the top tier are Claude Opus 4.6 (sixth) and GPT-5.5 (ninth).

SWE-Bench and Agent Arena: Same Direction, Different Scale

On SWE-bench Verified, Fable 5 recorded 95.0% — the highest published score on that benchmark. The translation to Agent Arena is not guaranteed or proportional. What the current data shows is that Fable 5’s benchmark ceiling and its real-world multi-turn performance are both pointing in the same direction, which is not a given. Opus 4.7’s SWE-bench numbers did not produce a proportional Agent Arena lead over GPT-5.4.

With 14,177 votes and ±2.10% confidence, the 12.94% reading is meaningful but not yet fully settled. As its count approaches the 25,000-30,000 range common for established models, the score will either hold or compress toward the field. The sub-dimension data — particularly error recovery — provides the structural explanation for why it might hold.