GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

OpenAI's Jalapeño Chip Goes Official: Broadcom-Designed, 50% Cheaper Than Nvidia, Ships End-2026

OpenAI has operated entirely on rented compute since its founding: Nvidia GPUs through Microsoft Azure, then expanded across CoreWeave, Oracle Cloud, and SpaceX’s Colossus cluster. On June 24, 2026, that dependency began its structured exit.

OpenAI and Broadcom jointly announced Jalapeño — OpenAI’s first custom-designed AI chip, internally codenamed Titan. Broadcom CEO Hock Tan reported early testing showing roughly 50% better cost efficiency than standard AI GPUs, with performance he described as “on par with Nvidia’s.” Systems integration is handled by Celestica.

What Jalapeño Does

The chip is inference-only. It is not a training accelerator. That distinction matters because OpenAI’s cost structure has two distinct compute buckets: model training (infrequent, high-intensity) and inference (continuous, scales with every ChatGPT request and API call). Jalapeño targets the latter — the ongoing operational cost that grows with user growth.

At ChatGPT’s scale and the API’s enterprise load, a 50% inference cost efficiency improvement is a structural margin change. The chip does not help OpenAI train GPT-6 or its successors; it makes running existing models cheaper at volume.

Timeline

Jalapeño deployments begin end-2026. Full commercial scale arrives H1 2028 — a 12-to-18-month ramp before the chip meaningfully shifts OpenAI’s cost base. Nvidia H100, H200, and Blackwell clusters cover inference in the interim and continue handling training.

The Custom Silicon Pattern

Every major AI lab has now made the same calculation:

LabCustom ChipPrimary Use
GoogleIronwood (TPU v8)Training + inference
AmazonTrainium / InferentiaTraining / inference
MetaMTIAInference (ads + AI)
AppleNeural EngineOn-device inference
OpenAIJalapeñoInference

At frontier inference volumes, custom silicon justifies the design investment. The economics are consistent: you rent until the volume makes owning cheaper, then you build. OpenAI is the last major frontier lab to reach that threshold publicly.

Broadcom’s position is worth noting. The company designs chips for Google (TPU line), Apple (Silicon co-design), and now OpenAI. At the center of multiple frontier-lab silicon programs, Broadcom’s role in custom AI chip development is now structurally significant and expanding.

Nvidia’s Exposure

Jalapeño is inference-only, and OpenAI continues depending on Nvidia for training. The threat to Nvidia from this announcement is incremental: directional, not immediate, and bounded to one workload type at one customer. Google’s TPUs have run production inference for years without displacing Nvidia from training or from non-Google hyperscalers. The pattern repeats.

What Jalapeño signals at the industry level is that custom silicon for inference has become standard operating procedure for frontier labs — not a competitive differentiator but a cost-of-scale requirement that every lab eventually builds toward.

Key Numbers

  • Chip name: Jalapeño (codenamed Titan)
  • Announced: June 24, 2026
  • Partners: Broadcom (design), Celestica (systems)
  • Workload: LLM inference only
  • Cost efficiency claim: 50% better than standard AI GPUs (Broadcom CEO)
  • Performance claim: “on par with Nvidia’s” (Broadcom CEO)
  • Initial deployment: end-2026
  • Full commercial scale: H1 2028