GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Kimi K3 Hits AA-Briefcase Elo 1543 — Second to Fable 5, at $10.57 Per Task and 56 Minutes Each

Artificial Analysis has scored Kimi K3 on AA-Briefcase, its private benchmark for agentic knowledge work. Kimi K3 lands at Elo 1543 — second in the field, behind only Claude Fable 5 (1574) and above every other frontier model tested.

AA-Briefcase tests models on realistic multi-step deliverables: spreadsheets, presentations, and UI mockups produced from thousands of complex input files. Tasks are drawn from data science, product management, banking operations, and industrial strategy. Performance rolls up into a single Elo based on correctness, analytical quality, and presentation quality.

The Leaderboard

ModelAA-Briefcase Elo
Claude Fable 51574
Kimi K31543
GPT-5.6 Sol (max)1501
Claude Sonnet 5 (max)1388
Claude Opus 4.8 (max)1347
Kimi K2.6816

The jump from K2.6 to K3 is +727 Elo points in one generation. That is the largest single-generation improvement AA has measured on this benchmark.

Where Kimi K3 Leads and Where It Falls Short

On analytical quality Kimi K3 scores Elo 1754 — actually slightly above Fable 5 (1744). On raw rubric correctness it passes 51% of tasks (Fable 5: 56%).

The weak dimension is presentation. Kimi K3’s presentation Elo is 1471, behind GPT-5.6 Sol max (1660) and Claude Opus 4.8 max (1492). It produces analytically strong outputs that require more formatting work downstream.

The Cost and Time Penalty

This is the headline caveat. Kimi K3 averages $10.57 per AA-Briefcase task — roughly 10x more than K2.6, and above Claude Opus 4.8.

The cost driver is architecture and pricing: K3 uses 83 turns per task on average (vs 67 for Fable 5 and 50 for GPT-5.6 Sol), generates heavy output token volumes, and runs at Moonshot’s listed rate of $3/$15 per million input/output tokens with a 90% discount only on cached tokens. Output tokens carry the pricing, and K3 produces a lot of them per task.

Time follows the same pattern: 56.4 minutes per task on average. The Kimi API’s first-party throughput is lower than competing providers, which compounds the turn-count problem.

For batch agentic workloads where quality matters more than cost or latency, K3 is now the second-best available option. For production deployments where economics govern, GPT-5.6 Sol max runs faster, costs less, and posts better presentation scores — even though it trails K3 on analytical depth.

Benchmark Note

AA-Briefcase is a private dataset. No contamination concerns from prior training runs. AA has not published task-level breakdowns beyond the Elo and dimensional scores.