GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Alibaba's Zhenwu M890 Claims 3x the Nvidia H20 — and a Silicon Roadmap Through 2028

Alibaba has launched the Zhenwu M890, its most capable AI chip to date, developed by its semiconductor subsidiary T-Head. The announcement came alongside a new LLM, a rack-scale server, and a published roadmap extending to 2028 — the clearest signal yet that Alibaba is treating silicon as a permanent strategic vertical, not a stopgap against Nvidia import restrictions.

Specifications

The M890 targets Nvidia’s H20, the only Hopper-generation chip currently legal to export to China. Alibaba claims roughly 3x the H20’s performance and 3x the compute of its own predecessor, the Zhenwu 810E.

SpecZhenwu 810EZhenwu M890
Memory96GB144GB HBM3
Interchip bandwidth700 GB/s800 GB/s
Computebaseline~3x 810E
Precision supportFP16, FP8FP32, FP16, FP8, FP4

The M890 delivers approximately 0.6 PFLOPs FP16, which puts it in the A100 class of compute density. Independent benchmarking has not yet corroborated the H20 performance claim; full specification disclosure from Alibaba remains incomplete.

Rack-Scale Deployment

Alongside the M890, T-Head announced the Panjiu AL128 server, a single-rack configuration integrating 128 M890 chips. At 0.6 PFLOPs per unit, the AL128 represents roughly 76.8 PFLOPs FP16 per rack — a meaningful density for enterprise deployments that want to avoid the complexity of multi-rack GPU clusters.

The M890 is purpose-built for agentic AI workloads. Its memory architecture is designed for the sustained context retention and multi-step reasoning chains that autonomous agents require, rather than the high-throughput batch inference optimization common in Western GPU designs.

Commercial Footprint

T-Head has shipped approximately 560,000 Zhenwu chips across 400 customers in 20 industries, making it the most commercially deployed domestic AI chip platform in China by a wide margin. The scale distinguishes Alibaba’s position from Huawei’s Ascend — which has broader government backing — and from smaller domestic players still in pilot stages.

Infrastructure Commitment and Roadmap

Alibaba has committed 380 billion yuan (approximately $53 billion) to cloud and AI infrastructure across a three-year horizon, with the Zhenwu platform as the silicon foundation.

The published roadmap:

  • Zhenwu M890: Q2 2026 — shipping now
  • Zhenwu V900: Q3 2027 — another claimed 3x performance improvement, 216GB memory, 1200GB/s bandwidth
  • Zhenwu J900: Q3 2028 — architectural and performance upgrades unspecified

If the V900 lands on schedule, Alibaba will have produced three successive generational leaps within four years — a pace consistent with hyperscaler custom silicon programs at Google (TPU) and Microsoft (Maia).

Strategic Context

US export controls have progressively cut off China from advanced Nvidia silicon: H100 exports were blocked in 2022, followed by H800 and A800 in late 2023, and H20 restrictions are now under active policy discussion in Washington. Jensen Huang publicly warned in April that restricting H20s and forcing DeepSeek workloads onto Huawei Ascend would be “a horrible outcome for America,” acknowledging that domestic Chinese alternatives are closing the gap faster than previously expected.

The Zhenwu M890 is paired with the Qwen3.7-Max LLM, which was engineered for continuous agent operation up to 35 hours with 1,000+ tool calls — a workload profile that matches the M890’s memory architecture. The vertical integration of chip, server, cloud, and model puts Alibaba in a similar structural position to Google’s TPU-Gemini stack, just several years behind on raw compute performance.

Analysts note that Alibaba has not released independently verifiable compute performance numbers beyond the 3x H20 claim, and the M890’s FP16 figure suggests a meaningful gap remains versus Nvidia’s H100 at 2.0 PFLOPs FP16. The gap is closing, but it has not closed.