GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

DeepSeek V4 Pro and Flash Pricing Jumps Up to 12x on August 16 as Peak/Off-Peak Rates Replace Flat Billing

DeepSeek is replacing flat-rate billing for its V4 model family with a peak and off-peak structure, effective 16:00 UTC on August 16. The change arrives alongside the official GA of DeepSeek-V4-Pro-0813, which had been in preview since April.

The New Price Table

ModelPeriodCache-Hit InputCache-Miss InputOutput
V4-FlashCurrent$0.0028$0.14$0.28
V4-FlashOff-peak$0.007$0.22$0.66
V4-FlashPeak$0.014$0.44$1.32
V4-ProCurrent$0.003625$0.435$0.87
V4-ProOff-peak$0.022$0.66$1.98
V4-ProPeak$0.044$1.32$3.96

Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC — seven hours per day, concentrated in the morning and overnight shift in Beijing. The remaining 17 hours of each day are off-peak.

What the Increases Actually Look Like

The largest single jump is V4-Pro’s cache-hit input rate at peak, which goes from $0.003625 to $0.044 per million tokens, a factor of 12.1x. That is the figure behind the “up to 1,100%” headline circulating after the announcement.

Output token rates tell a more consistent story. V4-Flash peak output rises from $0.28 to $1.32 — a 371% increase from current flat pricing, or 2.4x from off-peak to flat current, depending on which baseline you use. V4-Pro peak output rises from $0.87 to $3.96, a 355% increase.

Off-peak rates are still above current flat pricing across the board. V4-Pro off-peak output ($1.98) is 2.3x the current rate. For most Western deployments, which fall almost entirely outside the UTC 01:00–10:00 peak window, the effective increase is the off-peak column, not the peak one.

DeepSeek frames the structure as resource allocation: shift flexible workloads to cheaper hours rather than pricing out price-sensitive users.

V4-Pro’s GA Benchmarks

The pricing change launches alongside V4-Pro-0813 going GA after four months in preview. Independent evaluation from Artificial Analysis puts the model at 53 on the Intelligence Index, one point above V4-Flash 0731, which scores 52.

The gap is narrow. V4-Pro’s strongest result is GPQA Diamond at 93%, tying Claude Opus 5 and Kimi K3. It trails on coding: Terminal-Bench 2.1 at 79% is ten points behind Claude Opus 5’s 89%, and SciCode comes in at 49%, the weakest result among models in Artificial Analysis’s comparison. On GDPval-AA v2, it scores 55%, ahead of GPT-5.6 Luna (53%) but well behind Claude Opus 5 (67%) and Grok 4.6 (62%).

At $0.06 per Intelligence Index task, V4-Pro remains on Artificial Analysis’s Pareto frontier — the point where the price curve steepens toward far more expensive frontier models. V4-Flash ($0.03/task) is also on that line, which means DeepSeek now holds two Pareto positions simultaneously, one at roughly mid-tier intelligence and one at slightly above it, both cheaper than any competitor at comparable scores.

The Structural Shift

For the past year, DeepSeek’s defining characteristic was aggressive, flat, no-time-of-day pricing that absorbed volatility and made cost planning predictable. V4-Flash’s July launch at $0.14 input / $0.28 output was widely cited as a signal that frontier-adjacent performance was becoming a commodity.

The August 16 structure reverses that framing. Time of day is now an economic variable. Workloads that can wait — batch inference, nightly pipelines, large-scale eval runs — can hold costs near current levels by scheduling outside peak hours. Workloads that cannot wait pay peak rates that are 2–4x the off-peak price, and 3–12x current flat rates depending on token type.

The company has not explained the pricing change beyond “allocate resources more reasonably.” V4-Flash demand spiked immediately after its July launch and has run at capacity since; the peak window maps precisely to business hours in China’s primary time zones.

V4-Pro’s final flat-rate price window closes at 16:00 UTC on August 16. From that point, the cheapest path to DeepSeek’s flagship model is an off-peak workload scheduled outside a seven-hour daily block.