Claude Opus 5 Sets ARC-AGI-3 Record at 30.2% — 4x the Previous Best
Claude Opus 5 posted 30.2% on ARC-AGI-3, the abstract visual reasoning benchmark designed to test genuine generalisation rather than pattern recall. The previous record was 7.8%, held by GPT-5.6 Sol (Max). Fable 5, Anthropic’s prior flagship, sits in the low single digits on the same benchmark.
ARC-AGI-3 tasks are novel by construction: participants must infer transformation rules from a small number of input-output grid pairs and apply them to unseen examples. They cannot be solved by recalling training data. The benchmark is widely regarded as one of the few remaining tests that cleanly separates capability from memorisation at the frontier.
The Score Gap
| Model | ARC-AGI-3 | SWE-bench Verified |
|---|---|---|
| Claude Opus 5 | 30.2% | 96.0% |
| GPT-5.6 Sol (Max) | 7.8% | ~86% |
| Claude Fable 5 | ~5% | 95.0% |
At 30.2%, Opus 5 is not close to solving ARC-AGI-3 — the benchmark ceiling is 85%+ for expert humans. But the gap over the prior best is the largest single-run jump on the benchmark since it launched.
Intelligence Index
Artificial Analysis placed Claude Opus 5 (Max) and Claude Opus 5 (xHigh) as the top two models on its Intelligence Index as of July 28, ahead of Claude Fable 5 and GPT-5.6 Sol. The Index aggregates performance across coding, reasoning, and knowledge benchmarks weighted by task difficulty. Opus 5 holds both top slots simultaneously — the first time a single lab has occupied the top two positions.
Benchmark Profile
Opus 5 also posts 43.3% on Frontier-Bench v0.1, ahead of Fable 5 (33.7%) and more than double Opus 4.8 (18.7%). On standard agentic evals: SWE-bench Verified 96.0%, SWE-bench Pro 79.2%, DeepSWE v1.1 68.8%. GDPval-AA v2 ELO 1861, ranked first.
The one category where Opus 5 explicitly trails the frontier: dual-use capability evaluations (cybersecurity, biological). Anthropic’s system card notes the model was specifically trained to rate below Fable 5 and Mythos on those dimensions. The high-risk agentic tier stays behind the Glasswing access program.
Cost Position
Opus 5 standard pricing is $5 input / $25 output per million tokens. On LiveBench, the cost per successful task is $0.487 — versus $1.573 for Fable 5 Max Effort and $0.589 for GPT-5.6 Sol Max Effort. Opus 5 is 70% cheaper per task than Fable 5 while posting a higher overall LiveBench score in the Agentic Coding category (61.3% vs 46.9%).
Opus 5 Fast, launched July 24 at $10/$50 per million, offers higher throughput at 2x the standard price — still one-third the cost of Opus 4.6 Fast at launch.
What It Means
The ARC-AGI-3 result is the most scrutiny-resistant number in today’s benchmark stack. It cannot be inflated by prompt engineering or dataset contamination. A 30.2% score does not signal general reasoning at human level — but it does suggest Opus 5 is doing something qualitatively different from prior frontier models on abstract rule inference. Whether that advantage transfers to real-world tasks is the open question the field is actively measuring.