GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Claude Opus 5 Sets ARC-AGI-3 Record at 30.2% — 4x the Previous Best

Claude Opus 5 posted 30.2% on ARC-AGI-3, the abstract visual reasoning benchmark designed to test genuine generalisation rather than pattern recall. The previous record was 7.8%, held by GPT-5.6 Sol (Max). Fable 5, Anthropic’s prior flagship, sits in the low single digits on the same benchmark.

ARC-AGI-3 tasks are novel by construction: participants must infer transformation rules from a small number of input-output grid pairs and apply them to unseen examples. They cannot be solved by recalling training data. The benchmark is widely regarded as one of the few remaining tests that cleanly separates capability from memorisation at the frontier.

The Score Gap

ModelARC-AGI-3SWE-bench Verified
Claude Opus 530.2%96.0%
GPT-5.6 Sol (Max)7.8%~86%
Claude Fable 5~5%95.0%

At 30.2%, Opus 5 is not close to solving ARC-AGI-3 — the benchmark ceiling is 85%+ for expert humans. But the gap over the prior best is the largest single-run jump on the benchmark since it launched.

Intelligence Index

Artificial Analysis placed Claude Opus 5 (Max) and Claude Opus 5 (xHigh) as the top two models on its Intelligence Index as of July 28, ahead of Claude Fable 5 and GPT-5.6 Sol. The Index aggregates performance across coding, reasoning, and knowledge benchmarks weighted by task difficulty. Opus 5 holds both top slots simultaneously — the first time a single lab has occupied the top two positions.

Benchmark Profile

Opus 5 also posts 43.3% on Frontier-Bench v0.1, ahead of Fable 5 (33.7%) and more than double Opus 4.8 (18.7%). On standard agentic evals: SWE-bench Verified 96.0%, SWE-bench Pro 79.2%, DeepSWE v1.1 68.8%. GDPval-AA v2 ELO 1861, ranked first.

The one category where Opus 5 explicitly trails the frontier: dual-use capability evaluations (cybersecurity, biological). Anthropic’s system card notes the model was specifically trained to rate below Fable 5 and Mythos on those dimensions. The high-risk agentic tier stays behind the Glasswing access program.

Cost Position

Opus 5 standard pricing is $5 input / $25 output per million tokens. On LiveBench, the cost per successful task is $0.487 — versus $1.573 for Fable 5 Max Effort and $0.589 for GPT-5.6 Sol Max Effort. Opus 5 is 70% cheaper per task than Fable 5 while posting a higher overall LiveBench score in the Agentic Coding category (61.3% vs 46.9%).

Opus 5 Fast, launched July 24 at $10/$50 per million, offers higher throughput at 2x the standard price — still one-third the cost of Opus 4.6 Fast at launch.

What It Means

The ARC-AGI-3 result is the most scrutiny-resistant number in today’s benchmark stack. It cannot be inflated by prompt engineering or dataset contamination. A 30.2% score does not signal general reasoning at human level — but it does suggest Opus 5 is doing something qualitatively different from prior frontier models on abstract rule inference. Whether that advantage transfers to real-world tasks is the open question the field is actively measuring.