DeepSeek V4 Is Weeks Away: 1 Trillion Parameters, Huawei Chips, Apache 2.0
DeepSeek is expected to release V4 before the end of April. The original target was February. Then March. Then early April. The latest window — late April — has the strongest sourcing yet: Reuters, citing The Information, reported “within the next few weeks” on April 4. Multiple independent Chinese outlets including Sina Finance, iFeng, and IT之家 corroborated a late-April date following what they described as internal communications from founder Liang Wenfeng.
V4-Lite, a smaller variant, has been live on API infrastructure since early April. When DeepSeek stress-tests infrastructure, a full release typically follows within weeks.
The Specs (Confirmed vs Unverified)
| Claim | Source | Status |
|---|---|---|
| Huawei Ascend 910C/950PR chips | Reuters, TrendForce | Confirmed |
| ~1 trillion total parameters (MoE) | Multiple leaks | Unverified |
| ~37 billion active per token | Multiple leaks | Unverified |
| 1 million token context window | Widespread reports | Unverified |
| Apache 2.0 open weights | DeepSeek pattern + leaks | Likely |
| SWE-bench Verified ~81% | Pre-release testing | Unverified |
The Huawei chip detail is the one confirmed fact that matters most. Every prior DeepSeek release ran on NVIDIA H100s. V4 was trained entirely on Huawei hardware — a deliberate response to US export controls that cut off access to advanced NVIDIA GPUs. If V4 reaches frontier performance on non-NVIDIA silicon, it undercuts the core argument that export controls meaningfully slow Chinese AI development.
The Benchmark Claim
The 81% SWE-bench Verified figure cited in pre-release testing would put V4 ahead of Gemini 3.1 Pro (80.6%), ahead of the current GPT-5.4 (77.2%), and below Claude Opus 4.7 (87.6%) and Claude Mythos Preview (93.9%). At DeepSeek’s typical pricing — V3.2 runs at $0.27/M input — an 81% SWE-bench result at that cost would represent the most efficient frontier-class coding model on the market by a wide margin.
None of these numbers should be treated as final. DeepSeek publishes detailed technical reports with every major release — V3’s was 70 pages with full architectural details and training recipes. Until that paper drops, every spec is provisional.
The Architecture
V4 uses a Mixture-of-Experts design: roughly 1 trillion total parameters, with approximately 37 billion active per token. That ratio — 37B active out of 1T total — keeps inference costs comparable to a dense 37B model despite the full-parameter scale for specialised tasks. The same architectural pattern drove V3’s cost advantage.
The reported 1 million token native context window would be a significant jump from V3.2’s 128K. DeepSeek published research on Engram — a conditional memory system for long-context retrieval — in January 2026. V4-Lite in early testing shows materially improved context recall, which supports the claim.
What’s At Stake
DeepSeek has done this twice already. V3 shipped December 2024 and reset cost expectations across the industry. R1 followed in January 2025 and showed that chain-of-thought reasoning at frontier level was achievable in open weights. V4 completing that arc — on Huawei hardware, at scale, under Apache 2.0 — would be the third shock in 18 months.
Whether late April actually holds is the only open question. The infrastructure signal from V4-Lite suggests the infrastructure is ready. The rest is a launch date.