GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Claude Opus 4.8 Leads Intelligence Index at 61.4: AA Analysis Puts GDPval Elo at 1890, 67% Win Rate Over GPT-5.5

Artificial Analysis has published its independent analysis of Claude Opus 4.8, providing the first third-party benchmark breakdown since Anthropic’s launch announcement. The headline number is an Intelligence Index score of 61.4 — +4.1 points over Opus 4.7, and +1.2 points ahead of GPT-5.5 (xhigh), the previous leader.

Intelligence Index

Opus 4.8 scored 61.4 on the AA Intelligence Index. The field:

ModelIntelligence Index
Claude Opus 4.8 (max)61.4
GPT-5.5 (xhigh)60.2
GPT-5.5 (high)~58
Claude Opus 4.7 (max)57.3
Gemini 3.1 Pro Preview~56

The 4.1-point gain from 4.7 to 4.8 is the largest single-version increment in the Claude Opus line since AA began tracking.

GDPval-AA: the agentic performance number that matters

On GDPval-AA — AA’s primary evaluation for agentic performance on knowledge work tasks — Opus 4.8 scored 1,890 Elo at launch with its max effort setting. That is +137 points over Opus 4.7 and +121 points ahead of the second-ranked model, GPT-5.5 xhigh.

In head-to-head GDPval comparisons, a 121-point Elo gap implies approximately a 67% win rate against GPT-5.5 xhigh. Anthropic gets that lead while using 35% fewer output tokens than Opus 4.7 and completing tasks in 15% fewer turns. The one area where efficiency still trails: Opus 4.8 uses approximately 30% more turns than GPT-5.5 to complete equivalent GDPval tasks.

Scientific reasoning

Previous Opus releases have trailed peers on complex academic benchmarks. Opus 4.8 changes that positioning.

On Humanity’s Last Exam (HLE), Opus 4.8 leads by 1 point in what AA describes as a “tight contest” between Anthropic, Google DeepMind, and OpenAI. On CritPt — a frontier physics evaluation developed by Argonne National Laboratory and UIUC — Claude Opus 4.8 has moved ahead of Gemini 3.1 Pro, though it still trails GPT-5.4 and GPT-5.5 on that specific benchmark.

Pricing and Fast Mode

Standard pricing is $5/$25 per million tokens (input/output) — identical to Opus 4.7. Fast Mode is priced at 2x the standard rate ($10/$50), a reduction from the 6x premium that applied to Opus 4.7 Fast Mode. At current pricing, Opus 4.8 Fast Mode costs less per token than Opus 4.6 standard.

What the numbers mean for model selection

The GDPval gap is the more operationally significant figure. Intelligence Index scores measure general capability across a wide benchmark spread; GDPval specifically tests agentic performance on knowledge work at realistic task lengths. A 67% Elo-implied win rate on that task class — delivered with 35% fewer tokens — makes Opus 4.8 the most cost-efficient option in the agentic frontier tier on a per-correct-task basis, at least until GPT-5.5 xhigh is evaluated on GDPval under an equivalent effort setting.

SWE-bench Verified data for Opus 4.8 has not been independently published. Anthropic’s launch materials do not include the number; independent evaluation is pending. AA notes that Opus 4.8 improves across benchmarks relative to 4.7 without providing SWE-specific figures.