GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

DeepSeek Is Building Its Own Inference Chip to Break From Nvidia

DeepSeek is developing its own AI inference chip, according to multiple reports published July 7. The chip is designed for inference — the compute phase where a trained model generates responses for users — rather than training, where the engineering challenge is different. Nvidia shares fell in pre-market trading on the news.

The move marks a strategic shift for a company that built its reputation on extracting maximum capability from scarce, restricted GPU hardware. DeepSeek has operated under US export controls limiting access to Nvidia’s most capable chips since early 2024. Its efficiency-first engineering culture — which produced the algorithmic optimisations behind V3 and V4’s cost-performance ratios — now appears to be extending downstream into silicon design.

Why Inference, Not Training

Training a frontier model requires hundreds to thousands of tightly coupled GPUs operating in synchrony across weeks. It is an infrastructure problem that favours Nvidia’s NVLink interconnect and the software ecosystem built around CUDA over decades. DeepSeek has navigated that constraint by running training on Huawei Ascend clusters and whatever Nvidia H800 inventory it accumulated before restrictions tightened.

Inference is a different problem. Serving millions of user requests requires throughput, memory bandwidth, and latency optimisation at scale — but not the tight synchronisation of a training run. It is the same problem that pushed Google toward TPUs, Amazon toward Trainium, and Meta toward MTIA: the workload is predictable enough to design custom silicon around, and the unit economics of serving at scale justify the upfront chip investment.

DeepSeek’s DSpark framework, published in late June, demonstrated the lab’s approach: a speculative decoding architecture that boosted user-facing generation speed 60-85% over a baseline system while running on existing hardware. A custom inference chip is the hardware layer that DSpark was built to eventually run on.

The China Silicon Pattern

DeepSeek joining the inference chip race adds a new actor to a dynamic that has been developing for 18 months. Huawei’s Ascend 910B is already a functional alternative for training in China; Alibaba’s Zhenwu M890 targets 3x the Nvidia H20 on inference workloads. DeepSeek has so far been a model lab running on other people’s hardware. Building inference silicon is the move toward vertical integration.

The timing is not accidental. Export controls on Nvidia H20s — the chip that had been sold into China — tightened in April 2025. DeepSeek’s ability to expand inference capacity depends either on acquiring approved hardware through official channels (limited), sourcing on the grey market (expensive and legally exposed), or building their own path (capital-intensive but strategically durable). A proprietary inference chip resolves the dependency permanently.

Nvidia Context

Nvidia’s data center revenue hit $75.25B in Q1 FY27, up 92%. But the customer concentration risk is real: roughly half of that revenue comes from hyperscalers who are simultaneously funding Amazon Trainium, Google TPU, Microsoft Maia, and Meta MTIA — all designed to reduce Nvidia GPU spend per inference workload. DeepSeek’s chip effort is smaller in scale but larger in symbolism: a Chinese lab that became one of the most-used AI providers in the world is now engineering off the Nvidia dependency rather than lobbying around it.

Nvidia’s Kyber rack-scale architecture — intended to house Rubin Ultra chips — was separately reported this week to be delayed from 2027 to 2028, adding to investor caution. The pre-market pressure from the DeepSeek chip report landed on a stock already underperforming the broader semiconductor rally.

No specifications, partners, or timelines for DeepSeek’s chip effort have been disclosed. The reports cite multiple sources familiar with the plans.