Huawei Atlas 960E SuperPoD: 8 EFLOPS at FP8 Across 4096 NPUs, Replacing 48,000 Optical Modules With 5,500
Huawei launched the Atlas 960E SuperPoD at HUAWEI CONNECT 2026 in Shanghai on September 17, with a keynote from David Wang, Deputy Chairman and Rotating Chairman. The central claim: an NPO-based (Near-Packaged Optics) interconnect architecture that Huawei says is an industry first, enabling 4,096 NPUs to operate under unified memory addressing across a single SuperPoD.
The Numbers
At FP8 precision: 8 EFLOPS. At FP4: 16 EFLOPS. Those figures sit at the cluster level, not the chip level — the Atlas 960E is a full SuperPoD, not a new NPU.
The interconnect story is where the architecture diverges from conventional designs. Traditional SuperPoDs at this NPU count require approximately 48,000 800G optical modules to connect the compute fabric. The Atlas 960E uses 5,500 Hi-ONE units instead. Hi-ONE is Huawei’s NPO optical engine — a mass-produced unit delivering 7.2 Tbit/s transmission capacity with a built-in light source, which eliminates external light sources as a separate reliability bottleneck.
The power delta: 550 kW lower than optical-module equivalents. Availability: 99.8%, achieved by doubling the system’s fault-free operating time through tighter optical integration.
The Architecture
UnifiedBus, Huawei’s interconnect protocol, provides unified memory addressing across all 4,096 NPUs in the pod. This is the enabler for the NPO architecture: the entire cluster operates as a single-tier memory namespace rather than requiring multi-hop data movement across independent nodes.
The design uses an orthogonal architecture with fully liquid cooling. Huawei describes it as targeting training and inference for 10-trillion-parameter models — a scale that has largely defined the roadmap for next-generation frontier systems.
Adjacent Announcements
Two additional products launched alongside the Atlas 960E:
TaiShan 950 SuperPoD (upgraded): General-purpose compute rather than AI accelerators. Powered by the same UnifiedBus all-optical networking, supporting up to 4,096 nodes with a unified memory pool of up to 256 TB. The announced applications include higher-density agent sandboxes, faster agent sandbox startup times, and improved vector search performance for agent pipelines.
OceanStor M900: A UnifiedBus-powered context memory storage cluster for the L3.5 memory tier. Supports direct one-hop access and provides petabyte-scale KV cache capacity — targeting the inference side of large agentic deployments where context retrieval latency is a bottleneck.
CANN Goes Open Source
Wang announced that CANN — Huawei’s Compute Architecture for Neural Networks, the software layer that sits between AI frameworks and Ascend hardware — is now fully open source and in community-driven development. The move brings the Ascend ecosystem into direct comparison with CUDA’s developer base, with Huawei framing it as moving from “usable” to “user-friendly.”
Context
The Atlas 960E enters a market where Huawei is operating under sustained US chip export controls that restrict its access to advanced NVIDIA and TSMC-fabbed silicon. The NPO optical interconnect approach represents a system-level response: building cluster efficiency through interconnect density and protocol integration rather than competing on individual chip performance metrics.
HUAWEI CONNECT 2026 runs September 17-19 at the Shanghai World Expo Exhibition and Convention Center.