GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Anthropic Launches Claude Opus 4.6 Fast Mode at $30/$150 Per Million — 6x the Standard Price

Anthropic has added a fast-mode variant of Claude Opus 4.6, available now via API and OpenRouter. The capabilities are identical to the standard Opus 4.6 — the difference is entirely speed and price.

The Numbers

VariantInputOutputThroughputTTFT
Claude Opus 4.6$5/M$25/M~30 tok/s~2.5s
Claude Opus 4.6 Fast$30/M$150/M92 tok/s0.93s

The throughput delta is roughly 3x; the price delta is 6x. Cache reads on Fast Mode price at $3/M; cache writes at $37.50/M. Context window stays at 1M tokens, max output at 128K.

What It’s For

Anthropic’s documentation frames Fast Mode around latency-sensitive deployments: real-time voice pipelines, high-turn interactive agents, and user-facing applications where response lag is a product problem. For background processing, batch jobs, or workflows that run off the critical path, the cost premium is hard to justify.

The 6:1 price ratio suggests this is not intended as the default path. It’s an escape valve for the segment of production traffic where Opus-level capability and sub-second first-token latency must coexist.

Context

Anthropic’s standard Opus 4.6 is already the top composite performer on the Stack Futures ticker at 80.8% SWE-bench Verified and 99.3% tau-bench telecom. The Fast variant doesn’t change those figures — it just changes when you can afford to wait.

At $150/M output, Claude Opus 4.6 Fast Mode is the most expensive production API on the market. OpenAI’s GPT-5.4 (xhigh) runs cheaper per token for equivalent task scores. Whether that gap is worth paying depends entirely on latency requirements.

The model ID on OpenRouter is anthropic/claude-opus-4.6-fast.