Microsoft Cuts Image Generation Costs 41% With MAI-Image-2-Efficient — $19.50/M Output Tokens
Microsoft has launched MAI-Image-2-Efficient, a production-targeted image generation model that cuts output costs by 41% compared to its flagship MAI-Image-2. It shipped April 14 on Microsoft Foundry and MAI Playground with no waitlist.
Key Numbers
- Image output: $19.50/M tokens (down from $33/M on MAI-Image-2)
- Text input: $5/M tokens (unchanged)
- Speed: 22% faster than MAI-Image-2
- GPU efficiency: 4x higher throughput per NVIDIA H100 at 1024×1024
- Latency vs. Google: 40% faster at p50 than Gemini 3.1 Flash, Gemini 3.1 Flash Image, and Gemini 3 Pro Image (self-reported)
The Differentiation
Microsoft is positioning two distinct tiers. MAI-Image-2-Efficient is the volume workhorse — product shots, marketing assets, UI mockups, interactive workflows where fidelity tolerance is higher and throughput matters. MAI-Image-2 remains the precision tool for photorealistic portraits, intricate stylization, and work where output quality justifies the higher cost.
The framing reflects a market reality: text generation APIs have undergone 97% price compression over three years, but image generation has lagged. At $33/M output tokens, MAI-Image-2 sits above the batch-production threshold for many commercial buyers. At $19.50/M, the model enters competitive territory.
Context
MAI-Image-2 launched on Microsoft Foundry on April 2 alongside MAI-Transcribe-1 and MAI-Voice-1, as part of the MAI Superintelligence team’s push to build a full production AI platform independent of OpenAI integrations. Satya Nadella reorganized Copilot teams in March 2026 with explicit emphasis on cost reduction.
Shutterstock is listed as an early partner testing the Efficient variant. The model is rolling out to Copilot, Bing, and PowerPoint with additional surfaces expected ahead of Microsoft Build 2026.
The Azure catalog lists the model ID as MAI-Image-2e, version dated April 9, 2026. Context window is listed at 131,072 tokens.