GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

NVIDIA RTX Spark: 1 Petaflop, 6,144 Cores, 128GB — NVIDIA's First Personal AI Superchip

NVIDIA revealed RTX Spark at Computex 2026 in Taipei — a superchip that integrates a Blackwell GPU and a 20-core CPU into a single package for Windows laptops and small desktops. It is the first time NVIDIA has shipped its own combined processor for consumer PCs rather than selling discrete graphics cards.

The headline numbers: 1 petaflop of FP4 AI compute, up to 6,144 Blackwell GPU cores, up to 128GB of unified memory shared between CPU and GPU workloads, and what NVIDIA describes as the most power-efficient RTX chip ever made. The CUDA software stack runs natively, which matters for developers — any model that runs on an NVIDIA data centre GPU will run on RTX Spark without a port.

The Microsoft partnership is the commercial framing. NVIDIA is positioning this explicitly for personal AI agents on Windows, not primarily for gaming. The product page language: “Built for Agents and AI.” Gaming support is listed third, after agents and creative applications.

What Changed

Every prior RTX laptop product was a discrete GPU paired with a third-party CPU — Intel or AMD silicon running the machine, NVIDIA handling graphics. RTX Spark collapses that into one die with unified memory. It mirrors the architecture Apple adopted with M1 in 2020, where CPU and GPU share the same memory pool and eliminate the PCIe bandwidth bottleneck that limited GPU-to-CPU data movement.

The 128GB unified memory ceiling is the practical number that matters for local AI workloads. A 70B-parameter model in FP4 quantisation requires roughly 35GB. At 128GB, RTX Spark can hold two 70B models simultaneously, or run a 100B+ parameter model comfortably without offloading to storage.

Why It Matters

NVIDIA has spent five years building the data centre AI stack — H100, H200, Blackwell — while Apple quietly claimed the personal AI laptop market with M-series chips. RTX Spark is NVIDIA’s answer.

The CUDA moat is the strategic play. Enterprise developers write to CUDA. If a team’s production workload runs on H100 clusters and their local development machine also runs CUDA natively, the workflow unifies. The alternative — recompiling for Apple Silicon, dealing with Metal, or running CPU-only inference — breaks that loop.

At 1 petaflop FP4, RTX Spark sits roughly where the first-generation H100 PCIe (3.9 PFLOPS FP8) was for dense inference workloads, scaled to a laptop power envelope. That is not a data centre replacement, but it is enough compute to run real local inference against frontier-class quantised models at interactive latency.

NVIDIA has not announced pricing or OEM partners. “Notify me” on the product page suggests devices are not yet available. Given Computex timing and a typical 3-6 month ramp to retail, RTX Spark machines are most likely shipping Q3-Q4 2026.