NVIDIA Vera Rubin: 50 PFLOPS FP4, 3nm Chiplets, and a New GPU Built Purely for Long-Context Inference
NVIDIA’s Vera Rubin architecture is shipping to customers in the second half of 2026, and the gap over Blackwell is wider than the usual generational cadence. The headline number is 50 petaflops of FP4 inference per GPU — but the more significant change is a new product category with no predecessor: the Rubin CPX, a separate GPU designed specifically for long-context inference at scale.
The Process Shift
Blackwell used TSMC’s 4NP process. Vera Rubin moves to 3nm.
The transistor count goes from 208 billion to 336 billion — a 1.6x increase on a chip that maintains the same dual-die architecture introduced with Blackwell. Two reticle-sized chiplets connect at 10 TB/s die-to-die bandwidth. The same configuration, substantially more density.
NVIDIA also doubled the CPU-to-GPU interconnect again, continuing the bandwidth scaling that has defined each generation since Hopper. For distributed training and multi-GPU inference, interconnect throughput is often the binding constraint at scale — it determines how efficiently gradients sync and how seamlessly model weights can be sharded across chips.
Rubin CPX: The Long-Context Product
The Rubin CPX is the genuinely new element. It is a separate GPU from the standard Rubin, designed for long-context inference workloads — not training, not standard inference. No Blackwell chip played this role.
Long-context inference has different hardware requirements than standard inference. The attention mechanism’s memory and compute costs scale quadratically with sequence length, meaning a model processing a 200K-token context uses radically more memory bandwidth and on-chip memory than the same model at 8K tokens. The H100 and H200 can handle this, but they are general-purpose chips sized for training and standard inference — the cost per long-context token is high relative to the memory committed.
A chip purpose-built for long-context inference can optimize for high-bandwidth memory capacity and attention-specific compute patterns without carrying the training-optimized hardware that pushes Blackwell’s TDP to its current levels. If NVIDIA is building a standalone product for this, it signals that long-context inference is large enough at scale to justify bespoke silicon — which tracks with enterprise demand for RAG, legal document processing, and multi-step agentic tasks running on 100K+ token contexts.
NVL72: The Rack Configuration
The flagship server configuration is the NVL72 — 72 Rubin GPUs in a single liquid-cooled rack. Liquid cooling is required at this density; air cooling at this TDP and rack configuration is no longer viable for frontier inference or training hardware.
The 72-GPU rack configuration mirrors the approach used in Blackwell NVL72 systems but with higher per-chip compute. At 50 PFLOPS FP4 per GPU, an NVL72 rack delivers 3.6 exaflops of FP4 inference — an order of magnitude increase in rack-level throughput from what a comparable H100 system produced.
The Infrastructure Timeline
Vera Rubin ships to customers in H2 2026. For the AI infrastructure stack, that means the bulk of new Vera Rubin capacity will come online in 2027 — after procurement, power provisioning, facility buildout, and integration. The intermediate layer will run on Blackwell — particularly the NVL72 Blackwell systems already ordered by hyperscalers and AI labs — while Vera Rubin transitions into the next provisioning cycle.
Meta’s deal to take the bulk of CoreWeave’s Vera Rubin capacity through 2032, announced earlier this month, locks in priority allocation before other buyers can move on H2 supply. The infrastructure arms race rewards early commitment; by the time Vera Rubin is in general availability, the largest buyers will already have claimed most of the 2027 allocation.
The Rubin CPX timeline is less clear. NVIDIA has not published a ship date for the long-context product separate from the main Rubin SKU, but given that it involves a distinct die design and cooling requirements, it is likely a 2027 product even if Rubin GPU shipping begins in H2 2026.