CoreWeave First to Deploy NVIDIA Vera Rubin NVL72 at Rack Scale — 10x Inference Efficiency, MLPerf v6.0 Record
CoreWeave is the first AI cloud to bring up and fully validate NVIDIA’s Vera Rubin NVL72 at rack scale. The company completed system-level validation on June 1, 2026. Its stock closed at $120 that day, up 13.96%, recovering ground lost after a mixed Q1 earnings report three weeks earlier.
The Vera Rubin NVL72 is a rack-scale system: 72 Rubin GPUs and 36 Vera CPUs per rack, connected by sixth-generation NVLink fabric running 260 TB/s aggregate bandwidth. NVIDIA’s published efficiency claims against Blackwell: 10x better inference per watt, one-quarter the GPU count for equivalent throughput, one-tenth the cost per million tokens.
No other cloud provider has a production Vera Rubin deployment. AWS, Azure, and Google Cloud are on the allocation list but have not validated systems at rack scale yet. The moat is measured in quarters, not years — hyperscalers will receive their Rubin allocations — but first-to-market on a new generation translates directly into contract wins while rivals are still qualifying hardware.
MLPerf v6.0 Inference
Simultaneously, CoreWeave tops the MLPerf v6.0 inference benchmark results, posting 2x performance improvement over its prior submission. MLPerf is the primary industry benchmark for AI system performance; leading results signal that CoreWeave’s software optimization layer — not just the hardware — is keeping pace with each generation.
The company earned the top Platinum ranking in both SemiAnalysis ClusterMAX 1.0 and 2.0, and leads Artificial Analysis’s inference speed and price-performance rankings for Kimi K2.6 specifically. Moonshot AI’s K2.6 runs at 981 tokens per second on Cerebras for raw throughput, but CoreWeave holds the top position on inference speed benchmarks for GPU cloud deployments.
What Vera Rubin Changes
The efficiency improvements on Vera Rubin are not incremental. Moving from Blackwell to Rubin at the inference layer materially changes the economics of long-context workloads:
- Extended context inference (1M+ tokens) is where per-token cost currently limits practical deployment. One-tenth the cost per million tokens on the same workload changes the unit economics for any application running at scale.
- Agentic AI workflows requiring sustained throughput across multi-step task chains benefit directly from 10x inference per watt — it means the same power budget supports 10x the throughput, or the same throughput at one-tenth the power spend.
- The NVLink 6 fabric at 260 TB/s removes the inter-GPU communication bottleneck that currently forces model architects to shard large models conservatively. Trillion-parameter models become more tractable to serve.
The architecture uses Dell Technologies for server infrastructure and Micron liquid-cooled NVMe storage, indicating a purpose-built supply chain integration rather than off-the-shelf rack deployment.
Business Context
CoreWeave’s $88 billion revenue backlog is the central fact in its infrastructure thesis. The company went public in 2025 and posted Q1 revenue of $2.08 billion — double the year-ago quarter. The Q1 earnings miss (light Q2 guidance) came from a mismatch between contract value and recognized revenue, not from demand weakness. Vera Rubin deployment changes that conversation: hardware in production can be invoiced.
The two largest contracts feeding CoreWeave’s backlog are with Meta and Anthropic. Meta runs Llama 4 inference at scale across CoreWeave infrastructure. Anthropic’s Claude serves significant inference demand through AWS and CoreWeave GPU capacity. Both contracts predate Vera Rubin, but CoreWeave’s first-mover position on the new architecture gives it a renewal argument when those terms come up for renegotiation.
CoreWeave also announced its addition to the Nasdaq-100 index in June, the first AI infrastructure company to reach the index. The market is treating neocloud GPU capacity as a durable infrastructure asset rather than a cyclical buildout.
The Watch List
Q2 earnings (expected August) will show whether Vera Rubin utilization translates to revenue conversion. Two metrics will settle the debate: utilization rate across the Rubin fleet (above 80% validates the capacity bet; below 65% signals pricing pressure) and per-GPU revenue holding steady as AWS and Azure compete for the same enterprise inference accounts once they receive their own Rubin allocations.
The first Fortune 500 inference contract signed by a neocloud — outside the frontier labs themselves — will be the cleaner signal that the market has moved from lab-scale to enterprise-scale, and that CoreWeave has a durable seat at a table historically occupied by hyperscalers alone.