GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Meta Deploys MTIA 400: 6 Petaflops Custom Silicon Targets Nvidia for GenAI Inference

Meta’s MTIA 400 chip has cleared internal testing and is deploying across Meta data centers, according to reporting published this week. The deployment runs in parallel with Meta’s recently announced bulk purchases of Nvidia and AMD GPUs — a hedge that is starting to pay off on the inference side.

What MTIA Actually Does

MTIA stands for Meta Training and Inference Accelerator. It is not a general-purpose GPU and was never intended to be. Meta designed the chip to handle AI inference workloads inside its own infrastructure, not to be sold externally.

The MTIA 300, already widely deployed, handles Meta’s ranking and recommendation systems — the algorithms that select what appears in Facebook and Instagram feeds, processing billions of requests per day at inference-only workloads with no training component. That is the use case the original MTIA architecture was optimised for.

The MTIA 400 is a step-change in scope. The new target workloads are generative AI inference: image generation, video synthesis, and text-based AI responses. This is the workload that Nvidia’s H100 and H200 dominate commercially, and where Nvidia captures the highest margins.

MTIA 400 Specifications

  • FP8 compute: 6 petaflops
  • TDP: 1,200W
  • HBM capacity: 288GB
  • HBM bandwidth: 9.2 TB/s
  • Rack density: 72 chips per rack

For comparison, an H200 SXM5 delivers roughly 4 petaflops FP8 with 141GB HBM3e at 3.35 TB/s. MTIA 400 carries more than double the HBM capacity, which matters for serving large models without model parallelism overhead across chips.

The TDP is aggressive at 1,200W per chip, and rack density at 72 chips means power management at the facility level is a real constraint. But for Meta’s purpose — running a fixed set of workloads at known traffic patterns inside controlled data centers — the trade-offs are manageable in ways they would not be for a general-purpose cloud provider.

Why This Matters for Open-Source Economics

Meta’s MTIA deployment path is directly tied to the cost economics of its Llama API offerings. If Llama 4 Maverick and Scout can run on MTIA 400 infrastructure rather than leased Nvidia H200 capacity, Meta’s marginal cost of serving open-weight models drops substantially.

The Llama API, launched alongside Llama 4 in early 2026, competes on price against commercial providers. The price floor is set by inference hardware cost. Every MTIA 400 rack that can replace Nvidia GPU rack time for serving Llama workloads directly widens Meta’s margin or allows further price cuts — making open-weight models more commercially accessible from first-party infrastructure.

The Hedge Structure

Meta buying bulk Nvidia and AMD GPUs at the same time it deploys MTIA 400 is not contradictory. The GPU purchases cover training — MTIA is inference-only — and provide burst capacity for inference workloads that MTIA infrastructure has not yet scaled to cover. The hedge also insulates Meta from execution risk if MTIA deployment encounters reliability issues at scale.

The bet is that over a 3-5 year horizon, Meta can progressively replace Nvidia inference spending with MTIA silicon while keeping Nvidia dependency primarily on training, where no cost-effective alternative exists. That is the same architecture that Google pursued with TPUs, and which now underpins Google’s ability to offer Gemini API pricing below what most infrastructure math would support on third-party hardware.