AI API Pricing in 2026: A 97% Drop in Three Years and Still Falling
AI API pricing has undergone one of the fastest cost deflations in the history of enterprise software. When GPT-4 launched in March 2023, it cost $30 per million input tokens. Today, comparable-quality models run under $1/M — a 97% reduction in three years.
Current Pricing by Provider (April 2026)
Average blended cost across each provider’s lineup:
| Provider | Avg Input/1M | Models Tracked |
|---|---|---|
| Anthropic | $6.27 | 3 |
| OpenAI | $3.15 | 5 |
| Perplexity | $3.00 | 1 |
| Cohere | $2.50 | 1 |
| xAI | $1.65 | 2 |
| Mistral AI | $1.15 | 2 |
| Amazon | $0.80 | 1 |
| $0.68 | 2 | |
| DeepSeek | $0.41 | 2 |
| Alibaba Cloud | $0.22 | 2 |
| Meta | $0.17 | 2 |
The Spread That Matters
The 18x price gap between GPT-5.4 ($2.50/M input) and DeepSeek V3.2 ($0.14/M) on comparable workloads is the most consequential number in enterprise AI procurement right now. For a production app processing 10M tokens monthly, that’s $25 versus $1.40 — a difference that compounds fast at scale.
What’s Driving the Deflation
Three forces are compressing prices simultaneously:
- Open-source pressure — Meta Llama, DeepSeek, and Qwen have established a cost floor. Any proprietary model priced significantly above their capability-adjusted cost needs to justify the premium.
- Google’s aggressive positioning — Gemini Flash at $0.10-0.30/M sets the pace for fast inference tiers. Google can afford to price at or below cost on models to protect its cloud revenue.
- Continuous optimization — inference efficiency improvements (quantization, speculative decoding, batching) reduce serving costs quarter over quarter, and competitive pressure forces those savings to be passed on.
Where Anthropic Sits
Anthropic’s $6.27/M average is the highest in the industry — by design. Claude Opus 4.6 at $5/$25 is positioned as a premium reasoning model, not a commodity. The bet is that the performance gap justifies the price premium in high-stakes use cases (legal, medical, agentic). The SWE-bench and Tau2-bench numbers suggest that bet is currently paying off.
The question is how long the premium holds as Google, OpenAI, and open-source models close the capability gap at a fraction of the price.