Claude Opus 5 Takes AA-Briefcase #1 at Elo 1720 — 146 Points Clear of Fable 5, 20% Cheaper Per Task
Anthropic released Claude Opus 5 on July 24, 2026. Artificial Analysis benchmarked all five effort settings before release and published the AA-Briefcase results alongside the announcement. Opus 5 at max effort posts an AA-Briefcase Elo of 1720 — 146 points ahead of Claude Fable 5 (1574), which held the previous top position.
The margin matters: AA-Briefcase is not a saturating benchmark. The gap from second to first is larger than the gap from fifth to third in most of the models Artificial Analysis tracks.
AA-Briefcase Elo by Effort Tier
| Model | Elo | Cost per Task |
|---|---|---|
| Claude Opus 5 (max) | 1720 | $17.79 |
| Claude Opus 5 (xhigh) | 1693 | $14.26 |
| Claude Fable 5 (max) | 1574 | $22.30 |
| Claude Opus 5 (high) | 1606 | $10.41 |
Opus 5 at xhigh outperforms Fable 5 max effort by 119 Elo points at 64% of the cost. Opus 5 at high outperforms Fable 5 max by 32 points at 47% of the cost. There is no effort tier where Fable 5 matches or beats Opus 5 on the AA-Briefcase benchmark.
What AA-Briefcase Measures
AA-Briefcase is Artificial Analysis’s proprietary benchmark for agentic knowledge work. Tasks span 91 realistic private workloads requiring models to ingest thousands of input files and produce deliverables: research reports, presentations, and spreadsheets. Performance is scored across correctness, analytical quality, and presentation quality, combined into a single Elo.
The benchmark was specifically designed to test sustained capability over multi-step workflows rather than single-turn answer quality. Earlier this year AA-Briefcase revealed that Claude Opus 4.8 required 23 minutes per task and GPT-5.5 required 11 — operational throughput is part of the evaluation surface.
Agentic Coding Divergence
AA-Briefcase results diverge from LiveBench’s agentic coding leaderboard in a way that matters for buyers choosing between Opus 5 and Fable 5.
On LiveBench (as of July 25, 2026):
- GPT-5.6 Sol Max Effort leads overall at 82.4, including 65.6% agentic coding
- Claude Fable 5 Max Effort scores 80.8 overall but only 46.9% on agentic coding — the weakest score at the frontier tier
- Claude Opus 5 xHigh scores 80.3 overall and 61.3% on agentic coding — 14.4 points above Fable 5 on that dimension
Fable 5’s weak agentic coding score (46.9%) appears to reflect the same underlying gap AA-Briefcase exposes. Opus 5 closes that specific weakness while maintaining comparable overall intelligence.
Intelligence Index Position
Artificial Analysis also confirmed Opus 5 as the new leader on their Intelligence Index, displacing Fable 5. The Intelligence Index combines quality, reasoning, and coding benchmarks into a single composite. Having a model that leads both the Intelligence Index and the agentic knowledge work benchmark simultaneously is a new position for any lab — previously those rankings had split between Anthropic and OpenAI across different test surfaces.
Pricing
Opus 5 uses the same pricing tier as Opus 4.8. The per-task cost reduction on AA-Briefcase comes from efficiency improvements: Opus 5 completes the same 91-task evaluation in fewer tokens per task. The published per-million-token rate did not change.
Fable 5 remains available at $10/$50 per million tokens (credit-only as of July 1). Opus 5 enters at a lower per-task cost than Fable 5 on the benchmark most relevant to knowledge work — the use case both models are marketed for.