MiniMax M3 Tops Open-Weight Intelligence Index at 55 — Weights Still Missing on Day 8
Artificial Analysis published its independent evaluation of MiniMax M3 on June 8, placing it at 55 on the Intelligence Index — one point ahead of Kimi K2.6 (54) and MiMo-V2.5-Pro (54), and ahead of every Chinese open-weight peer.
The problem: MiniMax called M3 an “open-weight” model at its June 1 launch. Eight days later, no weights have appeared on Hugging Face. The company promised delivery “within approximately 10 days” of launch, targeting around June 11. That deadline arrives in two days.
What AA’s Evaluation Found
At Intelligence Index 55, M3 sits just below the closed proprietary frontier — Kimi K2.6 at 54, MiMo-V2.5-Pro at 54, DeepSeek V4 Pro below both. M3 is the first open-weight model to post above 54 on AA’s main index since the index was introduced.
On GDPval-AA — a real-world knowledge-work productivity benchmark — M3 scores approximately 1670 Elo, placing it level with Claude Sonnet 4.6 under Adaptive Reasoning at maximum effort. Once the weights land, it will also be the leading open-weight model on GDPval-AA.
Speed: 43.9 tokens per second on AA’s serving infrastructure.
The Vendor Benchmark Picture
MiniMax reported the following at launch. None of these have been independently verified at full scale:
| Benchmark | MiniMax M3 | Kimi K2.6 | GPT-5.5 | Claude Opus 4.8 |
|---|---|---|---|---|
| SWE-bench Pro | 59.0% | 58.6% | 58.6% | 69.2% |
| Terminal-Bench 2.1 | 66.0% | 66.7% | 78.2% | 74.6% |
| OSWorld-Verified | 70.1% | — | — | — |
| Context window | 1M tokens | 1T params | — | 1M tokens |
The SWE-bench Pro figure places M3 one point ahead of both Kimi K2.6 and GPT-5.5 on that benchmark — a first for an open-weight model against the closed frontier. However, these numbers are vendor-run, not independent. When MiniMax ran its M2.7 benchmarks, subsequent independent testing confirmed them within a few points, which is a reasonable prior.
The Open-Weight Problem
“Open-weight without weights” is an unusual launch posture. Developers cannot:
- Run the model on private infrastructure
- Fine-tune it for specific domains
- Verify the architectural claims in the technical report (also not published at launch)
- Know the license terms for commercial use
MiniMax’s previous model, M2.7, launched with weights and a commercially restricted license. If M3 ships the same structure, self-hosting teams will need to evaluate whether the license permits their use case before deploying.
The deadline is June 11. If weights do not arrive, the “open-weight” positioning becomes a genuine credibility issue for the company. If they do ship, M3 likely becomes the default recommendation for cost-sensitive long-context agentic work on non-sensitive data.
Competitive Position
The practical significance of Intelligence Index 55: closed proprietary models top out around 57-61 on the same index (GPT-5.5 at 60, Claude Opus 4.8 at 61). A one-point open-weight gap from Kimi K2.6 is not large, but the multimodal input and 1M-token context window differentiate M3 from K2.6’s text-and-code focus.
For organizations running agentic workloads that require video or image input at 1M-token context, M3 is the first open-weight option at frontier-adjacent intelligence. The question on June 11 is whether the weights confirm the architecture claims and whether the license is commercially usable.