Opus 4.7 Opens a 79-Point Agentic Elo Lead Over GPT-5.4 — and Runs Cheaper Per Task Than Its Predecessor
Claude Opus 4.7’s headline benchmark result is not the Intelligence Index tie at 57. It is the GDPval-AA Elo score: 1,753, measured across 44 occupations and 9 major industries in Artificial Analysis’s primary real-world agentic benchmark.
The gap is significant. The next closest models — Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) and GPT-5.4 (xhigh) — both score 1,674. Opus 4.7’s lead is 79 Elo points. Its predecessor, Opus 4.6, sits at approximately 1,619 — 134 Elo behind. On the benchmark that directly measures whether a model can execute complex multi-step knowledge work, Opus 4.7 is not merely ahead; it is in a different performance band.
The Cost Surprise
Opus 4.7 is priced at $5/$25 per million input/output tokens — identical to Opus 4.6 and Opus 4.5 at base pricing. But list price is not effective cost.
Opus 4.7 ships with a new tokenizer that reduces output token usage relative to Opus 4.6. Artificial Analysis estimates Opus 4.7 runs meaningfully cheaper per task than Opus 4.6 with Adaptive Reasoning and Max Effort enabled — which previously cost approximately $4,970 per task-equivalent run on AA’s benchmark suite. Opus 4.7 at max effort costs less while scoring 4 Intelligence Index points higher.
The practical implication: buyers who were running Opus 4.6 at full reasoning capacity and managing costs will find Opus 4.7 a strict improvement on both dimensions. Better outputs, lower per-task spend.
Hallucination Reduction
AA flags measurably lower hallucination rates for Opus 4.7 compared to Opus 4.6 at equivalent effort settings. Opus 4.7 ranks #2 on the AA Omniscience Index, behind Gemini 3.1 Pro Preview, but ahead of GPT-5.4 and its own predecessor. For enterprise use cases where factual reliability matters alongside task completion — legal, financial, medical knowledge work — the combination of GDPval-AA leadership and lower hallucination is a meaningful differentiator.
Key Numbers
| Metric | Opus 4.7 | Opus 4.6 | Sonnet 4.6 | GPT-5.4 |
|---|---|---|---|---|
| GDPval-AA Elo | 1,753 | ~1,619 | 1,674 | 1,674 |
| Intelligence Index | 57 | 53 | — | 57 |
| Pricing (input/output) | $5/$25 | $5/$25 | $3/$15 | $10/$40 |
| Context window | 1M tokens | 1M tokens | 200K | 128K |
Context
Opus 4.7 launched April 17, 2026. It maintains the 1M token context window from Opus 4.6. Anthropic has made several API changes alongside the release; specific details are documented in the API changelog. Opus 4.7 is distinct from Claude Mythos, which remains in limited preview access with reported 93.9% SWE-bench Verified performance.