Kimi K3 Leads Harvey LAB-AA at 26.7%, Nearly Double Fable 5 — Then Fixed 15 Bugs Fable Refused
Two data points published this week together establish something that individual benchmarks rarely do: Kimi K3, Moonshot AI’s 2.8-trillion-parameter open-weight model, is measurably more capable than Claude Fable 5 on legal work — and willing to do security work that proprietary models refuse.
Harvey LAB-AA: 26.7% vs 14.2%
Artificial Analysis’s Harvey LAB-AA benchmark runs 120 private legal tasks across 24 practice areas, including client memos, deposition summaries, and legal analysis documents. Each task uses a strict pass/fail rubric: every required item must be met for the task to count as solved. A single missed detail fails the entire assignment.
The current leaderboard:
| Model | Harvey LAB-AA |
|---|---|
| Kimi K3 | 26.7% |
| Claude Fable 5 | 14.2% |
Kimi K3’s score is 88% higher in absolute terms. In a benchmark where Fable 5 — one of the two strongest models on most leaderboards — succeeds on only 14 of 100 tasks, Kimi K3 succeeds on nearly 27. The strict grading explains why absolute numbers are low across the board: legal documents have no partial credit.
The benchmark was built by Artificial Analysis in partnership with Harvey, the legal AI company, and covers tasks drawn from real law firm workflows. It is private — models cannot be trained against it.
The Guardrails Gap
Separately, a user tasked Kimi K3, OpenAI Codex, and Claude Fable 5 with patching 15 critical security vulnerabilities across a production codebase. The job took 10 hours and cost $250 using Kimi K3’s API. Codex and Fable refused the work, citing cyber guardrail restrictions.
All 15 patches were completed by Kimi K3. Former White House AI czar David Sacks highlighted the result publicly, framing it as a policy question about where safety guardrails end and legitimate security work begins.
Kimi K3 is open-weight under a permissive license. The model can be run without any external safety layer. Frontier proprietary models increasingly apply restrictions at the inference layer that exclude categories of security work — vulnerability analysis, exploit proof-of-concepts, and patch generation for known CVEs — regardless of the legitimacy of the request. Whether that tradeoff is correctly calibrated is a live debate.
Context: Open-Weight Catching Up Fast
Rohan Paul’s AI newsletter, citing recent AISI research, notes that leading open-weight models now trail the closed frontier on long-horizon cyber capability by 4–7 months, down from 6–10 months through most of 2025. The gap is compressing by roughly 3 months per year.
Kimi K3’s SWE-bench Verified score of 93.4% places it third globally, behind Anthropic’s Mythos 5 (95.5%) and Fable 5 (95.0%) — the first open-weight model to enter that tier. On the AA Intelligence Index, it sits at #4 overall.
Demand has outrun supply: Moonshot AI suspended new Kimi K3 API subscriptions this week as compute capacity filled. US-based inference providers including Modal, Fireworks, and Baseten are expected to serve the model at roughly one-tenth the cost of Chinese providers, once Nvidia and AMD hardware allocation is confirmed.
The Composite Picture
The legal benchmark result and the guardrail story are different phenomena. One measures task completion on professional legal work. The other measures willingness to perform security operations that proprietary models decline.
Together they suggest that the practical capability ceiling for closed frontier models isn’t just a function of raw intelligence — it’s also a function of how broadly safety restrictions are applied. For workloads that sit in the overlap between “useful professional task” and “restricted category,” open-weight models with configurable guardrails now have a demonstrable advantage.