GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Samsung's zHBM Stacks DRAM Directly on AI Accelerators, Claims 8x HBM5 Performance

Samsung unveiled two next-generation 3D memory architectures at the Future of Memory and Storage (FMS) 2026 conference in California on August 5, with zHBM positioned as the most significant structural departure from conventional HBM design in a decade.

What zHBM Is

Conventional HBM places memory stacks horizontally around the AI accelerator die, connected via a silicon interposer. The layout limits how close memory can get to compute, which limits bandwidth and wastes power on long data paths.

zHBM eliminates the interposer entirely. Memory is stacked vertically on top of the AI accelerator along the z-axis, so the distance data travels between storage and processing collapses to a fraction of what current designs require. Samsung’s VP of DRAM Development Kim Kyung-ryun presented the figures:

  • 8x higher data-processing performance versus eighth-generation HBM5
  • 3x better performance per watt
  • 50%+ reduction in thermal resistance, improving system stability at sustained AI workloads

The architecture includes a customizable interlayer between the memory stack and accelerator die. Customers can embed IP blocks in this interlayer — expanding capacity for memory-bound workloads, or adding dedicated acceleration logic for specific inference patterns. That makes zHBM a system-in-package platform rather than a drop-in memory component.

zNAND-O and 400-Layer V10 BV-NAND

Samsung also introduced zNAND-O, a next-generation NAND flash family that applies Through-Silicon Via (TSV) technology to vertical NAND stacking — the same interconnect approach that makes HBM high-bandwidth. The architecture is structurally similar to SK Hynix and Sandisk’s High Bandwidth Flash (HBF), which signals that TSV-based flash is moving from prototype to competitive battleground.

The “O” in zNAND-O signals its primary target: on-device AI, where large model weights must be stored locally and accessed with low latency.

Alongside the zNAND-O prototype, Samsung revealed V10 BV-NAND with more than 400 layers. BV-NAND (Bonding Vertical NAND) separates cell array and peripheral circuitry onto different wafers before bonding them — enabling denser, faster storage in the same footprint. The V10 generation delivers approximately 58% higher memory density versus V9, with read, write, and I/O performance improvements across the board.

Why This Matters for AI Infrastructure

The HBM constraint has been one of the binding limits on GPU and AI accelerator scaling. As models grow and batch sizes increase, memory bandwidth — not raw compute — becomes the bottleneck during inference. Suppliers like SK Hynix have locked in HBM capacity commitments through 2033, and AMD’s MI400 ships with 432GB of HBM4 as a direct response to that pressure.

zHBM attacks the problem from a different angle: instead of adding more HBM around the die, it moves memory onto the die, compressing the data path. If Samsung can deliver the 8x figure in production at acceptable yield, it changes the economics of the AI accelerator memory hierarchy and reduces the advantage that goes to whoever locked in the most HBM supply.

Samsung said it plans close co-optimization work with customers to tailor zHBM configurations to specific AI processor architectures. No production timeline was disclosed. The technology was shown at FMS as mockups and prototypes; volume roadmap details are expected closer to the HBM5E and HBM6 transition window.