GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Claude Opus 4.8 Launches: Super-Agent Benchmark Leader, 84% on Mind2Web, Fast Mode 3x Cheaper

Anthropic released Claude Opus 4.8 today — same price as Opus 4.7, improved across every benchmark category it has been tested on. The upgrade ships alongside Dynamic Workflows in Claude Code, effort-level controls on claude.ai, and a threefold reduction in the cost of Fast Mode.

Benchmark results

Anthropic has not yet released a full SWE-bench Verified score for Opus 4.8, but early-access partner data establishes its standing in several key categories:

Super-Agent benchmark: Claude Opus 4.8 is the only model to complete every case end-to-end. GPT-5.5 does not. Prior Opus models do not. The benchmark covers translation, deep research, slide-building, and analysis agent workflows — and Anthropic describes the result as delivery at “parity on cost.”

Online-Mind2Web (browser agent / computer use): 84%, described as “a meaningful jump over both Opus 4.7 and GPT-5.5.” This is the most specific competitive number in the launch. Mind2Web is a web-based task completion benchmark; an 84% score at this scale puts Opus 4.8 at the top of the computer-use category.

CursorBench: Opus 4.8 exceeds Opus 4.7 at every effort level. Tool call efficiency is up — “fewer steps for the same intelligence,” which matters significantly in long agent loops where token waste compounds.

Legal Agent Benchmark: First model to break 10% on the all-pass standard. The headline sounds modest, but in legal AI evaluation — where tasks require end-to-end completion with no partial credit — 10% represents a generational gap over prior performance. The evaluation covers substantive legal research and review.

Pricing and Fast Mode

Opus 4.8 is priced identically to Opus 4.7: $15 per million input tokens, $75 per million output tokens (standard). Fast Mode — 2.5x speed — is now three times cheaper than it was for previous Opus models. The prior fast pricing was $30/$150 (Opus 4.6 Fast Mode launch price); the implied new Fast Mode pricing represents a meaningful cost reduction for latency-sensitive agentic workloads that previously found Opus too expensive to run at speed.

Effort controls on claude.ai

Users on claude.ai can now set the effort level Claude applies to a task — essentially controlling the compute budget per query. This is distinct from Dynamic Workflows (a separate feature in Claude Code) and applies to standard claude.ai interactions. Anthropic does not detail the mechanism, but the positioning matches OpenAI’s “thinking” slider approach, allowing cost-sensitive users to dial down and power users to dial up.

What changed under the hood

Anthropic’s description is: “more reliable and sharper in its judgement when performing agentic tasks.” The specific behaviors cited by early access partners:

  • Asks better clarifying questions before acting
  • Catches its own mistakes mid-task
  • Pushes back when a plan is unsound
  • Builds confidence checks before large multi-service changes
  • Carries context and style direction across long sessions

These are qualitative descriptions of improved long-context coherence and self-monitoring — exactly the failure modes that make current-generation models unreliable for unattended agent deployment.

Key Numbers

  • Pricing: $15/$75 per million tokens (same as Opus 4.7)
  • Fast Mode: 3x cheaper than previous Opus fast pricing
  • Fast Mode speed: 2.5x standard
  • Online-Mind2Web: 84% (top of computer-use / browser-agent category)
  • Super-Agent benchmark: 100% case completion (only model)
  • Legal Agent Benchmark: first to exceed 10% all-pass standard