GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Microsoft Plans Maia 300: 30% Better Tokens-Per-Dollar, 300,000 TSMC Chips Ordered for 2027

Microsoft is planning a fall announcement for Maia 300, its next-generation custom AI accelerator, and has been in discussions with TSMC over capacity for 300,000 chips to be delivered in 2027. The company internally claims the chip delivers more than 30% better tokens per dollar than the current silicon mix running in its Azure fleet.

The sourcing comes from Reuters and The Information, both citing people familiar with the discussions. Microsoft did not confirm the report.

Why This Matters

Microsoft runs a large mix of NVIDIA hardware for AI inference — H100s, H200s, and now GB200s at scale. Maia 300 is its mechanism to reduce that dependency and gain margin control on the inference layer. A 30% token-efficiency gain, if it holds at production scale, means Azure can run the same workloads at meaningfully lower cost or offer lower prices while maintaining margin.

The arithmetic matters at Azure scale. Microsoft’s AI services revenue is growing faster than its hardware cost basis can shrink. Custom silicon that compresses the inference cost curve extends the margin window while NVIDIA’s next generation arrives.

The Maia line is not new. Maia 100 launched in 2023 and powers some Azure OpenAI inference. Maia 300 is a generational jump — built for the transformer inference workloads that dominate Azure’s AI load today, including large-context runs and multi-agent pipelines.

Anthropic as Target Customer

Reuters reports Microsoft has been discussing Anthropic as a prospective Maia 300 customer. Anthropic already runs on Azure under its multi-year agreement with Microsoft; the pitch would be to move Claude inference from third-party NVIDIA silicon to Microsoft’s own chips, potentially at lower cost to Anthropic and higher margin to Azure.

Anthropic is in a strong position to evaluate this: it has its own chip ambitions (job postings for custom silicon were filed earlier this year) and it runs on Google TPUs and AWS Trainium alongside Azure. Adding Maia 300 to the mix would give it another negotiating variable and cost benchmark.

Whether Anthropic would actually commit volume to Maia 300 before it has production performance data is a separate question. The revenue relationship between the companies is significant enough that this is a real discussion rather than a sales pitch.

Competitive Context

Google deployed TPU v8 at Cloud Next ‘26 and earns external TPU revenue for the first time. Amazon runs Trainium 3 for its own models and is expanding external access. Microsoft is the last of the three major hyperscalers to develop custom inference silicon at scale — Maia 300 closes that gap.

The TSMC capacity order for 300,000 chips signals manufacturing intent, not just a design exercise. At 2027 delivery timing, Maia 300 would be in production roughly 12 months after launch — consistent with a fall 2026 announcement and 2027 Azure GA.

Intel shares moved on the Reuters report, reflecting concern about TSMC capacity allocation and what custom hyperscaler silicon means for its own AI chip ambitions.