GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Claude Opus 5 Takes AA-Briefcase #1 at Elo 1720 — 146 Points Clear of Fable 5, 20% Cheaper Per Task

Anthropic released Claude Opus 5 on July 24, 2026. Artificial Analysis benchmarked all five effort settings before release and published the AA-Briefcase results alongside the announcement. Opus 5 at max effort posts an AA-Briefcase Elo of 1720 — 146 points ahead of Claude Fable 5 (1574), which held the previous top position.

The margin matters: AA-Briefcase is not a saturating benchmark. The gap from second to first is larger than the gap from fifth to third in most of the models Artificial Analysis tracks.

AA-Briefcase Elo by Effort Tier

ModelEloCost per Task
Claude Opus 5 (max)1720$17.79
Claude Opus 5 (xhigh)1693$14.26
Claude Fable 5 (max)1574$22.30
Claude Opus 5 (high)1606$10.41

Opus 5 at xhigh outperforms Fable 5 max effort by 119 Elo points at 64% of the cost. Opus 5 at high outperforms Fable 5 max by 32 points at 47% of the cost. There is no effort tier where Fable 5 matches or beats Opus 5 on the AA-Briefcase benchmark.

What AA-Briefcase Measures

AA-Briefcase is Artificial Analysis’s proprietary benchmark for agentic knowledge work. Tasks span 91 realistic private workloads requiring models to ingest thousands of input files and produce deliverables: research reports, presentations, and spreadsheets. Performance is scored across correctness, analytical quality, and presentation quality, combined into a single Elo.

The benchmark was specifically designed to test sustained capability over multi-step workflows rather than single-turn answer quality. Earlier this year AA-Briefcase revealed that Claude Opus 4.8 required 23 minutes per task and GPT-5.5 required 11 — operational throughput is part of the evaluation surface.

Agentic Coding Divergence

AA-Briefcase results diverge from LiveBench’s agentic coding leaderboard in a way that matters for buyers choosing between Opus 5 and Fable 5.

On LiveBench (as of July 25, 2026):

  • GPT-5.6 Sol Max Effort leads overall at 82.4, including 65.6% agentic coding
  • Claude Fable 5 Max Effort scores 80.8 overall but only 46.9% on agentic coding — the weakest score at the frontier tier
  • Claude Opus 5 xHigh scores 80.3 overall and 61.3% on agentic coding — 14.4 points above Fable 5 on that dimension

Fable 5’s weak agentic coding score (46.9%) appears to reflect the same underlying gap AA-Briefcase exposes. Opus 5 closes that specific weakness while maintaining comparable overall intelligence.

Intelligence Index Position

Artificial Analysis also confirmed Opus 5 as the new leader on their Intelligence Index, displacing Fable 5. The Intelligence Index combines quality, reasoning, and coding benchmarks into a single composite. Having a model that leads both the Intelligence Index and the agentic knowledge work benchmark simultaneously is a new position for any lab — previously those rankings had split between Anthropic and OpenAI across different test surfaces.

Pricing

Opus 5 uses the same pricing tier as Opus 4.8. The per-task cost reduction on AA-Briefcase comes from efficiency improvements: Opus 5 completes the same 91-task evaluation in fewer tokens per task. The published per-million-token rate did not change.

Fable 5 remains available at $10/$50 per million tokens (credit-only as of July 1). Opus 5 enters at a lower per-task cost than Fable 5 on the benchmark most relevant to knowledge work — the use case both models are marketed for.