GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

DeepSeek Confirms Significant API Price Hike — V4-Flash Demand Overwhelmed Its Proof-of-Concept Pricing

DeepSeek has officially announced that its API pricing will go up significantly. The announcement follows the V4-Flash 0731 checkpoint’s release last week, which landed at 82.7% on Terminal-Bench 2.1 and became the fastest-growing model ever by token usage on Ollama — prompting the platform to add capacity across the US and Europe to handle the surge.

The current pricing of $0.14 per million input tokens and $0.28 per million output tokens was set to prove a point: that near-frontier capability could be delivered at a fraction of Western frontier model costs. That point has been made. Serving the resulting traffic at proof-of-concept prices no longer works once the world concedes it.

What V4-Flash 0731 Delivered

DeepSeek’s V4-Flash 0731 arrived on July 31 with benchmark numbers that surprised the market:

  • Terminal-Bench 2.1: 82.7% — second among open-weight models, trailing only GLM-5.2 at the time
  • SWE-bench Verified: 79.0%
  • DeepSWE: 54.4%
  • NL2Repo: 54.2%
  • Intelligence Index (Artificial Analysis): 50

At $0.03 per Artificial Analysis task, the model was 105x cheaper than Claude Fable 5 per completed task. That price-performance ratio drove the Ollama adoption spike. Capacity additions followed within days.

The Timing Problem

DeepSeek announced the hike one week after V4-Flash 0731 shipped. That timing is awkward for a specific reason: the competitive landscape has shifted.

Meta’s Muse Spark 1.2 launched this week at $1.25/$4.25 per million tokens — still more expensive than current DeepSeek rates, but with comparable capability scores and a more complete agent toolchain. GPT-5.6 Luna, OpenAI’s mid-tier in the 5.6 family, has moved to pricing that now competes directly in the range where DeepSeek built its user base.

At current rates, DeepSeek retains the price advantage by a large margin. After the hike, the arbitrage that drove adoption may narrow significantly. The model that reshaped enterprise AI budget assumptions will be priced differently than the model that caused those assumptions to form.

The Broader Pattern

This is the second DeepSeek pricing move in two months. In late June, the company temporarily doubled V4 output prices during peak Beijing hours — framed as a capacity management measure — before restoring standard rates. The current announcement is different: an outright upward revision with no stated temporary framing.

DeepSeek’s $7 billion raise at a $50B valuation completed earlier this year, with an Inner Mongolia data center at hyperscaler scale (1GW) under construction. The combination of fresh capital, rising infrastructure costs, and a demand base that has outgrown the original pricing model creates the conditions for a structural repricing rather than an incremental adjustment.

The degree of the increase has not been disclosed.

Key Numbers

  • Current pricing: $0.14/M input, $0.28/M output
  • Cost per AA task at current rates: $0.03 (105x cheaper than Claude Fable 5)
  • Terminal-Bench 2.1: 82.7%
  • SWE-bench Verified: 79.0%
  • Intelligence Index: 50
  • Nearest competitors after hike: Meta Muse Spark ($1.25/$4.25/M), GPT-5.6 Luna