GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Opus 4.7 Opens a 79-Point Agentic Elo Lead Over GPT-5.4 — and Runs Cheaper Per Task Than Its Predecessor

Claude Opus 4.7’s headline benchmark result is not the Intelligence Index tie at 57. It is the GDPval-AA Elo score: 1,753, measured across 44 occupations and 9 major industries in Artificial Analysis’s primary real-world agentic benchmark.

The gap is significant. The next closest models — Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) and GPT-5.4 (xhigh) — both score 1,674. Opus 4.7’s lead is 79 Elo points. Its predecessor, Opus 4.6, sits at approximately 1,619 — 134 Elo behind. On the benchmark that directly measures whether a model can execute complex multi-step knowledge work, Opus 4.7 is not merely ahead; it is in a different performance band.

The Cost Surprise

Opus 4.7 is priced at $5/$25 per million input/output tokens — identical to Opus 4.6 and Opus 4.5 at base pricing. But list price is not effective cost.

Opus 4.7 ships with a new tokenizer that reduces output token usage relative to Opus 4.6. Artificial Analysis estimates Opus 4.7 runs meaningfully cheaper per task than Opus 4.6 with Adaptive Reasoning and Max Effort enabled — which previously cost approximately $4,970 per task-equivalent run on AA’s benchmark suite. Opus 4.7 at max effort costs less while scoring 4 Intelligence Index points higher.

The practical implication: buyers who were running Opus 4.6 at full reasoning capacity and managing costs will find Opus 4.7 a strict improvement on both dimensions. Better outputs, lower per-task spend.

Hallucination Reduction

AA flags measurably lower hallucination rates for Opus 4.7 compared to Opus 4.6 at equivalent effort settings. Opus 4.7 ranks #2 on the AA Omniscience Index, behind Gemini 3.1 Pro Preview, but ahead of GPT-5.4 and its own predecessor. For enterprise use cases where factual reliability matters alongside task completion — legal, financial, medical knowledge work — the combination of GDPval-AA leadership and lower hallucination is a meaningful differentiator.

Key Numbers

MetricOpus 4.7Opus 4.6Sonnet 4.6GPT-5.4
GDPval-AA Elo1,753~1,6191,6741,674
Intelligence Index5753—57
Pricing (input/output)$5/$25$5/$25$3/$15$10/$40
Context window1M tokens1M tokens200K128K

Context

Opus 4.7 launched April 17, 2026. It maintains the 1M token context window from Opus 4.6. Anthropic has made several API changes alongside the release; specific details are documented in the API changelog. Opus 4.7 is distinct from Claude Mythos, which remains in limited preview access with reported 93.9% SWE-bench Verified performance.