Groq Raises $650M to Rebuild as Inference Neocloud After Nvidia Took Its IP and Founders
Groq raised $650 million on May 28 from existing investors to fund a pivot into AI inference cloud services — five months after licensing its Language Processing Unit (LPU) architecture to Nvidia for a reported $20 billion and distributing $7.6 billion to shareholders.
The round is effectively guaranteed. Existing backers Disruptive and Infinitum have committed to backstop the full $650 million if other shareholders decline their pro-rata participation. Groq 2.0 is led by interim CEO Adam Winter and interim CFO Matt Eng. The founding team — CEO Jonathan Ross and president Sunny Madra — joined Nvidia as part of the December 2025 deal.
What Groq Was and What It Sold
Groq’s LPU architecture posted throughput numbers no GPU could match. Llama 4 Scout runs at over 460 tokens per second on GroqCloud. On an Nvidia H100, the same model delivers roughly 100 to 150 tokens per second. For Llama 3.3 70B, Groq achieves around 800 tokens per second against 120 to 140 on GPU alternatives.
Nvidia paid $20 billion for a perpetual, non-exclusive license to that architecture. The deal was structured as a license rather than an acquisition, which kept Groq as an independent legal entity and preserved its right to use its own IP — but it also handed Nvidia the same IP. The Groq 3 LPU appeared inside Nvidia’s Vera Rubin platform at GTC 2026 within months of the deal closing.
Senators Elizabeth Warren and Richard Blumenthal called it a “reverse acqui-hire” in March 2026 and opened a formal DOJ and FTC inquiry. That investigation remains open.
The Neocloud Bet
Groq’s pivot is straightforward: stop building chips for sale and start selling inference from those chips. GroqCloud is now positioned as an inference neocloud — a cloud provider differentiated by proprietary hardware optimized for a single workload rather than general-purpose compute.
The business inherits 3.5 million developers and a Fortune 500 customer pipeline built during the chip era. The $650 million will fund infrastructure expansion — more LPU clusters — and the engineering rebuild Groq needs after losing its founding team.
The structural vulnerability is clear. Nvidia already has the LPU architecture and the engineers who designed it. The Vera Rubin Ultra platform, expected in the second half of 2027, will combine Nvidia’s silicon with HBM4E memory and LPU inference acceleration, available through AWS, Azure, and Google Cloud at scale.
The 18-Month Window
When Nvidia begins offering LPU-speed inference through its existing cloud partnerships, Groq’s speed advantage becomes table stakes rather than a differentiator. The company has an estimated 18 to 24 months to build enough enterprise contracts, infrastructure scale, and developer lock-in to survive as a viable alternative.
Three signals will determine the outcome: whether Groq’s 3.5 million developer base holds through Q3 2026 when Groq 3 ships and Nvidia starts direct inference capacity sales; whether the antitrust inquiry accelerates or stalls; and whether other inference-focused hardware startups follow the same playbook of licensing IP rather than scaling as chip makers.
The $650 million is the fuel for that race. Whether the distance is enough is the question.