US and UK Safety Institutes Score Kimi K3 at 32.2% on ExploitBench — Stronger Than Any Prior Open-Weight Model
The US Commerce Department’s CAISI and the UK Artificial Intelligence Security Institute jointly published a preliminary assessment of Kimi K3’s offensive cyber capabilities on July 23. It is the first government-level evaluation of Moonshot AI’s open-weight model and confirms both the anxiety in Washington about Chinese open-weight competition and the meaningful gap that still exists.
ExploitBench Scores
| Model | ExploitBench | Notes |
|---|---|---|
| US frontier models (ceiling) | 76.2% | Safeguards disabled |
| Kimi K3 | 32.2% | Default configuration |
| GLM-5.2 | 24.4% | Default configuration |
The US frontier numbers are ceiling measurements, taken with safety constraints disabled, and are not representative of deployed product behavior. Kimi K3 and GLM-5.2 were assessed in standard configurations.
Multi-Step Attack Autonomy
In simulated multi-step corporate cyberattack scenarios:
- Strongest US frontier models reached step 28.5 on average
- Kimi K3 stopped at step 17 on average
The assessment’s language on Kimi K3: it “can autonomously execute meaningful portions of an attack, occasionally completes an entire simulated enterprise attack, and does not reliably refuse offensive requests.”
The institutes characterize Kimi K3 as “substantially behind frontier US models in offensive cyber capability” while simultaneously being “stronger than the previous leading open-weight model.” The previous benchmark holder is not named; GLM-5.2’s 24.4% ExploitBench score places it below Kimi K3 and establishes the prior ceiling for open-weight models.
Why the Comparison Is Asymmetric
The 44-point ExploitBench gap between US frontier models (76.2%) and Kimi K3 (32.2%) reflects different measurement conditions. US models were stripped of guardrails to establish capability ceilings. Kimi K3 was tested as it ships. A fairer ceiling-to-ceiling comparison would require running Kimi K3 without safety constraints — which the assessment did not do.
What the numbers do confirm: the open-weight tier now includes a model capable of autonomous attack execution, even if that capability sits far below the restricted-access US frontier.
Context: What Kimi K3 Is
Kimi K3 launched July 17, 2026. It is a 2.8 trillion parameter open-weight model from Beijing-based Moonshot AI, available commercially at $3 per million input tokens and $15 per million output tokens. On SWE-bench Verified it scores 93.4%, placing it third globally behind Anthropic’s Mythos 5 (95.5%) and Fable 5 (95.0%).
The cyber assessment is a narrow capability slice. CAISI and UK AISI both noted this is preliminary. A full capability evaluation across broader risk categories has not been published.