Claude Opus 4.8 Leads Intelligence Index at 61.4: AA Analysis Puts GDPval Elo at 1890, 67% Win Rate Over GPT-5.5
Artificial Analysis has published its independent analysis of Claude Opus 4.8, providing the first third-party benchmark breakdown since Anthropic’s launch announcement. The headline number is an Intelligence Index score of 61.4 — +4.1 points over Opus 4.7, and +1.2 points ahead of GPT-5.5 (xhigh), the previous leader.
Intelligence Index
Opus 4.8 scored 61.4 on the AA Intelligence Index. The field:
| Model | Intelligence Index |
|---|---|
| Claude Opus 4.8 (max) | 61.4 |
| GPT-5.5 (xhigh) | 60.2 |
| GPT-5.5 (high) | ~58 |
| Claude Opus 4.7 (max) | 57.3 |
| Gemini 3.1 Pro Preview | ~56 |
The 4.1-point gain from 4.7 to 4.8 is the largest single-version increment in the Claude Opus line since AA began tracking.
GDPval-AA: the agentic performance number that matters
On GDPval-AA — AA’s primary evaluation for agentic performance on knowledge work tasks — Opus 4.8 scored 1,890 Elo at launch with its max effort setting. That is +137 points over Opus 4.7 and +121 points ahead of the second-ranked model, GPT-5.5 xhigh.
In head-to-head GDPval comparisons, a 121-point Elo gap implies approximately a 67% win rate against GPT-5.5 xhigh. Anthropic gets that lead while using 35% fewer output tokens than Opus 4.7 and completing tasks in 15% fewer turns. The one area where efficiency still trails: Opus 4.8 uses approximately 30% more turns than GPT-5.5 to complete equivalent GDPval tasks.
Scientific reasoning
Previous Opus releases have trailed peers on complex academic benchmarks. Opus 4.8 changes that positioning.
On Humanity’s Last Exam (HLE), Opus 4.8 leads by 1 point in what AA describes as a “tight contest” between Anthropic, Google DeepMind, and OpenAI. On CritPt — a frontier physics evaluation developed by Argonne National Laboratory and UIUC — Claude Opus 4.8 has moved ahead of Gemini 3.1 Pro, though it still trails GPT-5.4 and GPT-5.5 on that specific benchmark.
Pricing and Fast Mode
Standard pricing is $5/$25 per million tokens (input/output) — identical to Opus 4.7. Fast Mode is priced at 2x the standard rate ($10/$50), a reduction from the 6x premium that applied to Opus 4.7 Fast Mode. At current pricing, Opus 4.8 Fast Mode costs less per token than Opus 4.6 standard.
What the numbers mean for model selection
The GDPval gap is the more operationally significant figure. Intelligence Index scores measure general capability across a wide benchmark spread; GDPval specifically tests agentic performance on knowledge work at realistic task lengths. A 67% Elo-implied win rate on that task class — delivered with 35% fewer tokens — makes Opus 4.8 the most cost-efficient option in the agentic frontier tier on a per-correct-task basis, at least until GPT-5.5 xhigh is evaluated on GDPval under an equivalent effort setting.
SWE-bench Verified data for Opus 4.8 has not been independently published. Anthropic’s launch materials do not include the number; independent evaluation is pending. AA notes that Opus 4.8 improves across benchmarks relative to 4.7 without providing SWE-specific figures.