GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Sonnet 5 Costs 15% More Per Task Than Opus 4.8, Despite Cheaper Per-Token Pricing

Artificial Analysis published its independent analysis of Claude Sonnet 5 on June 30, the same day Anthropic shipped the model, and the headline number is counterintuitive: at standard pricing, Sonnet 5 costs more per task than the Opus 4.8 it is supposed to undercut.

The AA Intelligence Index puts Sonnet 5 at 53 on max effort — the same score as GPT-5.5 with high reasoning, two to three points behind Opus 4.8 max and GPT-5.5 xhigh, and six points above Sonnet 4.6. That makes Sonnet 5 the fifth-ranked model globally by AA’s quality composite.

The Token Economics Problem

The per-token pricing looks straightforward: Sonnet 5 at $3 input / $15 output is 40% cheaper on output than Opus 4.8 ($5 input / $25 output). But the model doesn’t use the same number of tokens to complete tasks.

On max effort, Sonnet 5 uses approximately 40% more output tokens per AA Intelligence Index task than Sonnet 4.6. On the AA-Briefcase knowledge-work benchmark and GDPval-AA, it uses roughly 3x as many agentic turns as its predecessor. AA computed the fully-loaded cost per Intelligence Index task at standard pricing: Sonnet 5 runs $2.29, Opus 4.8 runs $1.99 — a 15% premium for the model that costs less per token.

This is the core tension in effort-based pricing. The effort dial is a capability dial, but it is also a spend dial. A model that defaults to high-effort problem-solving generates more output tokens than a model that satisfices. When pricing per token, the savings from a lower rate can be consumed entirely by a higher rate of token generation.

What the Promotional Window Changes

Anthropic is offering $2 input / $10 output through August 31, 2026 — a one-third reduction on both sides. At that rate, the per-task math inverts: Sonnet 5 becomes meaningfully cheaper than Opus 4.8 for most workloads. AA’s analysis uses standard pricing throughout, so the promotional discount is additive.

After September 1, the choice becomes more deliberate. Developers building agentic pipelines who need max-effort performance will need to decide whether Sonnet 5’s quality tier ($2.29/task at standard) justifies the premium over Opus 4.8 ($1.99/task), or whether they route to lower effort levels where Sonnet 5 is genuinely cheaper.

The effort system now has five levels on both Sonnet 5 and Opus 4.8: low, medium, high, xhigh, and max. Sonnet 5 adds xhigh to the prior four Sonnet 4.6 settings. AA’s data shows that performance scales meaningfully with effort on Sonnet 5 — the gap between low and max effort is roughly 6x in turn count on GDPval-AA.

Benchmark Numbers

On tasks where Sonnet 5 directly competes with Opus 4.8, the results are close:

  • Terminal-Bench 2.1: Sonnet 5 at 80.4%, Sonnet 4.6 at 67.0%, Opus 4.8 at 82.7%
  • SWE-bench Pro: Sonnet 5 at 63.2%, Sonnet 4.6 at 58.1%, Opus 4.8 at 69.2%
  • Humanity’s Last Exam (with tools): Sonnet 5 at 57.4%, Opus 4.8 at 57.9%
  • OSWorld-Verified: Sonnet 5 at 81.2%, Opus 4.8 at 83.4%
  • GDPval-AA v2: Sonnet 5 at 1,618, Opus 4.8 at 1,615

The GDPval-AA v2 result is the outlier: Sonnet 5 leads on agentic knowledge work despite being the smaller model. AA-Briefcase shows the same pattern. For reasoning-intensive tasks — the CritPt physics benchmark and comparable heavy-reasoning evals — Opus 4.8 maintains a meaningful gap, with Sonnet 5 trailing GLM-5.2, Opus 4.8, and Fable 5.

What This Changes for Developers

The launch article framed Sonnet 5 as a 7.5x output-cost saving versus Opus 4.8 at introductory pricing. That is accurate at the per-token level. It is less accurate at the per-task level, where the real cost depends on how many tokens the model generates per completed task — a number that varies significantly by effort setting, workload type, and model behavior.

AA’s analysis supports using Sonnet 5 for agentic knowledge work, where it matches or beats Opus 4.8. For reasoning-intensive benchmarks and coding at the frontier, Opus 4.8 and Fable 5 maintain a clear lead. The promotional window through August 31 makes Sonnet 5 the obvious choice for volume workloads regardless of task type. The calculus after September 1 is more nuanced.

Context window: 1 million tokens. Cache pricing: $3.75/M for cache writes (25% premium), $0.30/M for cache hits (90% discount). 5-minute TTL.