Xiaomi Cuts MiMo-V2.5 API Prices by Up to 86% Effective Today — Pro Drops to $0.87/M Output
Xiaomi repriced its entire MiMo-V2.5 API series at midnight CST on May 27, cutting output costs by 71–86% and resetting Token Plan quotas with a 5–8x capacity increase.
The new overseas API rates:
| Model | Input (old → new) | Output (old → new) | Output cut |
|---|---|---|---|
| MiMo-V2.5 | $0.40/M → $0.14/M | $2.00/M → $0.28/M | −86% |
| MiMo-V2.5-Pro | $1.00/M → $0.435/M | $3.00/M → $0.87/M | −71% |
Cache-hit pricing on inputs falls further: $0.0028/M for base, $0.0036/M for Pro — effectively free for heavy prompt-caching workflows.
What the Models Do
MiMo-V2.5-Pro is a 1T-parameter / 42B-active MoE with 1M context, benchmarked at SWE-bench Pro 57.2, Claw-Eval 63.8, and τ3-bench 72.9 — on par with Claude Opus 4.6 and GPT-5.4 across most agentic tasks. The model uses 40–60% fewer tokens per trajectory than frontier closed-source equivalents at those capability levels. A demonstrated autonomous task: built a complete SysY compiler in Rust (233/233 tests, 672 tool calls, 4.3 hours) without human intervention.
MiMo-V2.5 (base) is the omnimodal variant — native image, video, audio, and text understanding with the same 1M context — positioned for everyday coding tasks at half the token cost of Pro. On MiMo Coding Bench, it matches Pro on routine dev tasks.
Both models are already compatible with Claude Code, OpenCode, and Kilo as drop-in backends.
Context
Xiaomi has shipped three major model families since December 2025: V2-Flash, V2-Pro/Omni/TTS, and now V2.5. The company has committed $8.7B in AI investment through 2028. This pricing move follows DeepSeek’s permanent 75% V4-Pro cut and continues the structural compression of frontier API economics.
The V2 series is deprecated as of today. Xiaomi recommends migrating to V2.5 models.
Key Numbers
- MiMo-V2.5 output: $0.28/M (was $2.00)
- MiMo-V2.5-Pro output: $0.87/M (was $3.00)
- 1M context window: no additional cost multiplier
- SWE-bench Pro: 57.2 (Pro), comparable to Opus 4.6 and GPT-5.4
- Token efficiency: 40–60% fewer tokens per trajectory vs Claude Opus 4.6