Meta Deploys MTIA 400: 6 Petaflops Custom Silicon Targets Nvidia for GenAI Inference
Meta’s MTIA 400 chip has cleared internal testing and is deploying across Meta data centers, according to reporting published this week. The deployment runs in parallel with Meta’s recently announced bulk purchases of Nvidia and AMD GPUs — a hedge that is starting to pay off on the inference side.
What MTIA Actually Does
MTIA stands for Meta Training and Inference Accelerator. It is not a general-purpose GPU and was never intended to be. Meta designed the chip to handle AI inference workloads inside its own infrastructure, not to be sold externally.
The MTIA 300, already widely deployed, handles Meta’s ranking and recommendation systems — the algorithms that select what appears in Facebook and Instagram feeds, processing billions of requests per day at inference-only workloads with no training component. That is the use case the original MTIA architecture was optimised for.
The MTIA 400 is a step-change in scope. The new target workloads are generative AI inference: image generation, video synthesis, and text-based AI responses. This is the workload that Nvidia’s H100 and H200 dominate commercially, and where Nvidia captures the highest margins.
MTIA 400 Specifications
- FP8 compute: 6 petaflops
- TDP: 1,200W
- HBM capacity: 288GB
- HBM bandwidth: 9.2 TB/s
- Rack density: 72 chips per rack
For comparison, an H200 SXM5 delivers roughly 4 petaflops FP8 with 141GB HBM3e at 3.35 TB/s. MTIA 400 carries more than double the HBM capacity, which matters for serving large models without model parallelism overhead across chips.
The TDP is aggressive at 1,200W per chip, and rack density at 72 chips means power management at the facility level is a real constraint. But for Meta’s purpose — running a fixed set of workloads at known traffic patterns inside controlled data centers — the trade-offs are manageable in ways they would not be for a general-purpose cloud provider.
Why This Matters for Open-Source Economics
Meta’s MTIA deployment path is directly tied to the cost economics of its Llama API offerings. If Llama 4 Maverick and Scout can run on MTIA 400 infrastructure rather than leased Nvidia H200 capacity, Meta’s marginal cost of serving open-weight models drops substantially.
The Llama API, launched alongside Llama 4 in early 2026, competes on price against commercial providers. The price floor is set by inference hardware cost. Every MTIA 400 rack that can replace Nvidia GPU rack time for serving Llama workloads directly widens Meta’s margin or allows further price cuts — making open-weight models more commercially accessible from first-party infrastructure.
The Hedge Structure
Meta buying bulk Nvidia and AMD GPUs at the same time it deploys MTIA 400 is not contradictory. The GPU purchases cover training — MTIA is inference-only — and provide burst capacity for inference workloads that MTIA infrastructure has not yet scaled to cover. The hedge also insulates Meta from execution risk if MTIA deployment encounters reliability issues at scale.
The bet is that over a 3-5 year horizon, Meta can progressively replace Nvidia inference spending with MTIA silicon while keeping Nvidia dependency primarily on training, where no cost-effective alternative exists. That is the same architecture that Google pursued with TPUs, and which now underpins Google’s ability to offer Gemini API pricing below what most infrastructure math would support on third-party hardware.