GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Gimlet Labs Raises $300M at $3B to Route AI Inference Across Any Chip

Gimlet Labs has raised $300 million in a Series B led by Andreessen Horowitz, valuing the company at $3 billion. The round closes six months after an $80 million Series A, bringing total funding to $392 million.

The company’s pitch is routing AI inference across different chip architectures — not committing to any single hardware vendor. Gimlet calls itself the first multi-silicon inference cloud. The practical claim is that workloads can shift between Nvidia GPUs, Google TPUs, Amazon Trainium, and custom accelerators based on availability, latency, and cost.

Why the Valuation Jump Makes Sense

The $3B figure is a 37x step-up from Series A valuation in six months, which is the kind of number that invites skepticism. But the market context supports it.

Every hyperscaler is now building proprietary silicon specifically to reduce dependency on Nvidia’s rack hardware. Google has TPU v6. Amazon has Trainium 2. Microsoft has Maia 2. None of those chips runs the same stack. An operator who wants to serve production inference across all three faces a fragmented runtime problem — and that’s the problem Gimlet is selling against.

The Nvidia dynamic makes the abstraction layer more valuable, not less. Nvidia’s revenue is not at risk in the near term; custom silicon takes years to reach parity on arbitrary workloads. But for specific inference patterns — especially attention-heavy transformer decoding — custom accelerators are already price-competitive. An inference cloud that can dispatch to the cheapest matching substrate for a given model shape has a real efficiency advantage over one that is locked to a single chip family.

What $392M Buys

Gimlet is building infrastructure, not models. The capital goes toward:

  • Multi-datacenter footprint across chip types
  • Routing and scheduling software that handles dispatch and failover across heterogeneous hardware
  • Partnerships with hardware vendors who want inference demand for their non-Nvidia silicon

The a16z lead is notable. Ben Horowitz’s firm bet early on custom silicon (investments in Groq, Cerebras predecessors) and has been consistent in the view that the inference layer is where the durable margin in AI infrastructure will sit.

The Competitive Set

Gimlet is not competing with Together AI, Fireworks AI, or Baseten directly — those companies run inference on specific hardware. Gimlet’s positioning is the abstraction above them: a single API that dispatches to whatever compute is cheapest and available for the workload.

The closer comparisons are Crusoe Energy (which routes to stranded compute) and Anyscale (which manages distributed compute orchestration). Neither operates at the chip-abstraction level Gimlet is targeting.

Whether the routing layer captures lasting margin or gets commoditized by cloud providers building their own chip-agnostic APIs is the open question. The $3B bet is that it takes long enough to commoditize that Gimlet can establish defensible infrastructure and customer relationships first.