Kimi K3 Hits AA-Briefcase Elo 1543 — Second to Fable 5, at $10.57 Per Task and 56 Minutes Each
Artificial Analysis has scored Kimi K3 on AA-Briefcase, its private benchmark for agentic knowledge work. Kimi K3 lands at Elo 1543 — second in the field, behind only Claude Fable 5 (1574) and above every other frontier model tested.
AA-Briefcase tests models on realistic multi-step deliverables: spreadsheets, presentations, and UI mockups produced from thousands of complex input files. Tasks are drawn from data science, product management, banking operations, and industrial strategy. Performance rolls up into a single Elo based on correctness, analytical quality, and presentation quality.
The Leaderboard
| Model | AA-Briefcase Elo |
|---|---|
| Claude Fable 5 | 1574 |
| Kimi K3 | 1543 |
| GPT-5.6 Sol (max) | 1501 |
| Claude Sonnet 5 (max) | 1388 |
| Claude Opus 4.8 (max) | 1347 |
| Kimi K2.6 | 816 |
The jump from K2.6 to K3 is +727 Elo points in one generation. That is the largest single-generation improvement AA has measured on this benchmark.
Where Kimi K3 Leads and Where It Falls Short
On analytical quality Kimi K3 scores Elo 1754 — actually slightly above Fable 5 (1744). On raw rubric correctness it passes 51% of tasks (Fable 5: 56%).
The weak dimension is presentation. Kimi K3’s presentation Elo is 1471, behind GPT-5.6 Sol max (1660) and Claude Opus 4.8 max (1492). It produces analytically strong outputs that require more formatting work downstream.
The Cost and Time Penalty
This is the headline caveat. Kimi K3 averages $10.57 per AA-Briefcase task — roughly 10x more than K2.6, and above Claude Opus 4.8.
The cost driver is architecture and pricing: K3 uses 83 turns per task on average (vs 67 for Fable 5 and 50 for GPT-5.6 Sol), generates heavy output token volumes, and runs at Moonshot’s listed rate of $3/$15 per million input/output tokens with a 90% discount only on cached tokens. Output tokens carry the pricing, and K3 produces a lot of them per task.
Time follows the same pattern: 56.4 minutes per task on average. The Kimi API’s first-party throughput is lower than competing providers, which compounds the turn-count problem.
For batch agentic workloads where quality matters more than cost or latency, K3 is now the second-best available option. For production deployments where economics govern, GPT-5.6 Sol max runs faster, costs less, and posts better presentation scores — even though it trails K3 on analytical depth.
Benchmark Note
AA-Briefcase is a private dataset. No contamination concerns from prior training runs. AA has not published task-level breakdowns beyond the Elo and dimensional scores.