Samsung's zHBM Stacks DRAM Directly on AI Accelerators, Claims 8x HBM5 Performance
Samsung unveiled two next-generation 3D memory architectures at the Future of Memory and Storage (FMS) 2026 conference in California on August 5, with zHBM positioned as the most significant structural departure from conventional HBM design in a decade.
What zHBM Is
Conventional HBM places memory stacks horizontally around the AI accelerator die, connected via a silicon interposer. The layout limits how close memory can get to compute, which limits bandwidth and wastes power on long data paths.
zHBM eliminates the interposer entirely. Memory is stacked vertically on top of the AI accelerator along the z-axis, so the distance data travels between storage and processing collapses to a fraction of what current designs require. Samsung’s VP of DRAM Development Kim Kyung-ryun presented the figures:
- 8x higher data-processing performance versus eighth-generation HBM5
- 3x better performance per watt
- 50%+ reduction in thermal resistance, improving system stability at sustained AI workloads
The architecture includes a customizable interlayer between the memory stack and accelerator die. Customers can embed IP blocks in this interlayer — expanding capacity for memory-bound workloads, or adding dedicated acceleration logic for specific inference patterns. That makes zHBM a system-in-package platform rather than a drop-in memory component.
zNAND-O and 400-Layer V10 BV-NAND
Samsung also introduced zNAND-O, a next-generation NAND flash family that applies Through-Silicon Via (TSV) technology to vertical NAND stacking — the same interconnect approach that makes HBM high-bandwidth. The architecture is structurally similar to SK Hynix and Sandisk’s High Bandwidth Flash (HBF), which signals that TSV-based flash is moving from prototype to competitive battleground.
The “O” in zNAND-O signals its primary target: on-device AI, where large model weights must be stored locally and accessed with low latency.
Alongside the zNAND-O prototype, Samsung revealed V10 BV-NAND with more than 400 layers. BV-NAND (Bonding Vertical NAND) separates cell array and peripheral circuitry onto different wafers before bonding them — enabling denser, faster storage in the same footprint. The V10 generation delivers approximately 58% higher memory density versus V9, with read, write, and I/O performance improvements across the board.
Why This Matters for AI Infrastructure
The HBM constraint has been one of the binding limits on GPU and AI accelerator scaling. As models grow and batch sizes increase, memory bandwidth — not raw compute — becomes the bottleneck during inference. Suppliers like SK Hynix have locked in HBM capacity commitments through 2033, and AMD’s MI400 ships with 432GB of HBM4 as a direct response to that pressure.
zHBM attacks the problem from a different angle: instead of adding more HBM around the die, it moves memory onto the die, compressing the data path. If Samsung can deliver the 8x figure in production at acceptable yield, it changes the economics of the AI accelerator memory hierarchy and reduces the advantage that goes to whoever locked in the most HBM supply.
Samsung said it plans close co-optimization work with customers to tailor zHBM configurations to specific AI processor architectures. No production timeline was disclosed. The technology was shown at FMS as mockups and prototypes; volume roadmap details are expected closer to the HBM5E and HBM6 transition window.