Chinese AI Models Now Hold 30% of US Enterprise Token Use — and Peaked at 46%
Chinese AI models have crossed a threshold in US enterprise adoption that was not visible in aggregate download numbers: they now handle more than 30% of tokens consumed by US companies on OpenRouter every week, a figure that peaked at 46% and has not dropped back below the 30% floor since February 8.
The baseline for comparison: the 12-month average through mid-2025 was 11%. In the first half of 2025 it sat at 4.5%. The shift over the following 12 months is a compression of roughly a decade’s worth of enterprise software substitution into a single year.
The Cost Pressure Is the Mechanism
OpenRouter data quoted in CNBC puts Chinese open-weight models at 60-90% cheaper than leading US frontier alternatives. That spread did not always exist — the first generation of capable Chinese models was price-competitive but not benchmark-competitive. The current cohort, led by DeepSeek V4-Pro, GLM-5.2, and Hy3, is both.
“Price is doing the work here,” Harpreet Arora, head of agentic infrastructure at Vercel, told CNBC. “When a task doesn’t need the best model, teams are beginning to route it to the cheapest one that’s good enough, and the recent wave of models coming out of China is winning that trade.”
That routing logic is measurable in the Vercel data. GLM-5.2, released mid-June, saw daily token volume grow 27x and customer count grow 80x in its first full week on Vercel. That is the fastest adoption of any model tracked by Vercel in 2026.
Case Study: Lindy Moves 100% to DeepSeek
The clearest single data point is Lindy, a US AI productivity startup that moved its entire inference stack — 100% of traffic — off Anthropic’s Claude and onto DeepSeek. CEO Flo Crivello described the cost curve as going “down, like, crash to the ground,” saying the switch will save Lindy millions of dollars within months.
Lindy is not an edge case. Kyle Chan, a fellow at the Brookings Institution’s John L. Thornton China Center, told CNBC: “Chinese AI models are particularly attractive to American companies now as AI costs skyrocket. Where previously US companies were prioritizing AI adoption regardless of model, now they’re getting more cost-conscious.”
The Policy Accelerant
The adoption curve has a policy dimension the cost story alone does not capture. At the end of June, OpenAI restricted new model rollouts to trusted partners at the government’s request. Export controls on Anthropic’s Mythos and Fable models were enforced and then lifted after a 19-day standoff. Companies that needed production-ready inference during that period had to route around the access wall — and Chinese open-weight models were the available alternative.
DeepSeek is also benefiting from a separate mechanism: enterprises that cannot get Mythos-level US frontier access are finding that DeepSeek V4-Pro posts 80.6% on SWE-bench Verified for $0.87/M output — a benchmark score and a price that would have been the frontier ceiling 12 months ago.
What This Does Not Mean
The 30% figure is token share on OpenRouter — a developer platform weighted toward teams that route prompts across multiple providers. It is not representative of all enterprise AI spend. Companies locked into Azure, AWS, or Google Cloud for procurement reasons are not in this pool. OpenAI and Anthropic still dominate by revenue and likely by enterprise seat count.
What the data captures is the emerging layer of cost-optimized, model-agnostic routing infrastructure. That layer is growing faster than any individual vendor. And in that layer, Chinese models — open-weight, cheap, technically competitive — are winning 30-46% of the traffic.
Model Scores at Issue
For the models driving this shift, the benchmark gap with US frontier is shrinking:
| Model | SWE-bench Verified | Pricing (Output/M) |
|---|---|---|
| Claude Fable 5 | 95.0% | $75 |
| GPT-5.5 | 88.7% | ~$60 |
| GLM-5.2 | 84.2% | $3.50 |
| DeepSeek V4-Pro | 80.6% | $0.87 |
| Tencent Hy3 | 78.0% | OpenRouter free (2 weeks) |
The benchmark gap at the very top is still real. The pricing gap is 60-90x. For the majority of production tasks that do not require frontier-tier coding, the math is doing the routing work automatically.