GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Kimi K3 Takes Frontend Code Arena #1 at 1679 — First Chinese Model to Beat Fable 5 and GPT-5.6

Moonshot AI’s Kimi K3 landed at the top of Arena’s Frontend Code leaderboard on July 16, 2026, posting 1679 ELO across 1,757 blind developer votes. It beat Claude Fable 5 at 1631 and GPT-5.6 Sol at 1618 — the first time a Chinese model has topped that specific leaderboard.

The jump was not gradual. K2.6 sat at 18th place with 1515 ELO. K3 leaped 17 positions and 164 ELO points in a single generation.

What the numbers say

K3 ranked first in 6 of 7 frontend categories: Brand & Marketing, Reference-Based Design, Data & Analytics, and three others. Its only second-place finish was in the Game category. The leaderboard carries a preliminary tag given vote count maturity, but the signal is real: these are blind human developer evaluations on actual web development tasks, not static benchmarks.

The comparison to Fable 5 is stark. Anthropic’s current flagship — which holds the highest AA Intelligence Index score and 95% SWE-bench Verified — lost this specific evaluation by 48 ELO points. On overall intelligence rankings, Fable 5 still leads. On frontend code, K3 won.

Architecture

K3 is the largest open-weight model released to date: 2.8 trillion total parameters across 896 experts, with 16 experts active per token — roughly 1.8% of the pool per forward pass. Context window is 1 million tokens with native vision support.

Full weights are due on July 27. Until then, K3 is API-only.

Pricing

  • Cached input: $0.30/M tokens
  • Cache miss input: $3/M tokens
  • Output: $15/M tokens

That output price matches GPT-5.5 standard and sits above Fable 5’s API pricing. K2’s input cost was $0.60/M; uncached K3 input runs 5x higher, and output 25x higher than K2’s rate. The performance premium comes with a real cost step-up.

The context

This result matters beyond the leaderboard position. Frontend coding is where developers feel model quality immediately — UI, styling, component structure, interactivity. It is not a synthetic benchmark. It is preferences expressed by developers doing their actual jobs.

Moonshot acknowledges K3 still trails Fable 5 and GPT-5.6 Sol on overall intelligence. The company says K3 outperformed every other model in its evaluation suite including Claude Opus 4.8 and GPT-5.5 across coding and agentic tasks. That is the competitive tier K3 is claiming.

The political dimension was immediate. David Sacks commented publicly. The story hit Hacker News front page with substantial discussion. The reaction reflects a broader anxiety: open-weight Chinese models are no longer trailing frontier proprietary systems across the board.

Key numbers

ModelFrontend Code ELOOverall AA Index
Kimi K31679~58 (est.)
Claude Fable 5163164.9
GPT-5.6 Sol1618~63 (est.)

K3’s overall AA Intelligence Index rank sits below Fable 5 and GPT-5.6 Sol — confirmed by Moonshot itself. The frontend coding win is real but narrow: one leaderboard, one task domain, early vote counts. The weight release on July 27 will let the developer community run its own evaluations and determine whether K3’s Arena showing holds under independent testing.