GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

AI API Pricing in 2026: A 97% Drop in Three Years and Still Falling

AI API pricing has undergone one of the fastest cost deflations in the history of enterprise software. When GPT-4 launched in March 2023, it cost $30 per million input tokens. Today, comparable-quality models run under $1/M — a 97% reduction in three years.

Current Pricing by Provider (April 2026)

Average blended cost across each provider’s lineup:

ProviderAvg Input/1MModels Tracked
Anthropic$6.273
OpenAI$3.155
Perplexity$3.001
Cohere$2.501
xAI$1.652
Mistral AI$1.152
Amazon$0.801
Google$0.682
DeepSeek$0.412
Alibaba Cloud$0.222
Meta$0.172

The Spread That Matters

The 18x price gap between GPT-5.4 ($2.50/M input) and DeepSeek V3.2 ($0.14/M) on comparable workloads is the most consequential number in enterprise AI procurement right now. For a production app processing 10M tokens monthly, that’s $25 versus $1.40 — a difference that compounds fast at scale.

What’s Driving the Deflation

Three forces are compressing prices simultaneously:

  1. Open-source pressure — Meta Llama, DeepSeek, and Qwen have established a cost floor. Any proprietary model priced significantly above their capability-adjusted cost needs to justify the premium.
  2. Google’s aggressive positioning — Gemini Flash at $0.10-0.30/M sets the pace for fast inference tiers. Google can afford to price at or below cost on models to protect its cloud revenue.
  3. Continuous optimization — inference efficiency improvements (quantization, speculative decoding, batching) reduce serving costs quarter over quarter, and competitive pressure forces those savings to be passed on.

Where Anthropic Sits

Anthropic’s $6.27/M average is the highest in the industry — by design. Claude Opus 4.6 at $5/$25 is positioned as a premium reasoning model, not a commodity. The bet is that the performance gap justifies the price premium in high-stakes use cases (legal, medical, agentic). The SWE-bench and Tau2-bench numbers suggest that bet is currently paying off.

The question is how long the premium holds as Google, OpenAI, and open-source models close the capability gap at a fraction of the price.