GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Arena Opens Image-to-WebDev Leaderboard; Kimi K2.5 and Grok 4.20 Beta Enter Document Rankings

Chatbot Arena added a new evaluation surface on April 15: the Image-to-WebDev leaderboard, which judges models on their ability to convert images and design mockups into working web applications. The category puts visual-to-code capability — a task that matters directly for front-end development workflows — under the same human-preference methodology Arena applies to its text and code rankings.

The same week, four models joined Arena’s Document leaderboard: Moonshot AI’s Kimi K2.5 Thinking, xAI’s Grok 4.20 Beta (reasoning), Google’s Gemma 4 31B, and Anthropic’s Claude Opus 4.6 Thinking. The Document arena tests models on multi-page document comprehension, extraction, and synthesis — tasks closer to real legal, financial, and research workflows than standard chat benchmarks.

Kimi K2.5 Thinking: 1 Trillion Parameters, Open-Sourced

Kimi K2.5 is Moonshot AI’s most capable open model. Released January 27, it is a Mixture-of-Experts architecture with 1 trillion total parameters and 32 billion active per forward pass — the same efficiency-focused MoE design used by Mixtral and Grok. Moonshot pretrained it on approximately 15 trillion mixed visual and text tokens, making multimodality native rather than bolt-on.

Key specs:

  • Context window: 256K tokens
  • Modes: Instant, Thinking (chain-of-thought), Agent, and Agent Swarm
  • Architecture: 61-layer MoE, 400M-parameter MoonViT vision encoder
  • Availability: Open weights on HuggingFace; API on Moonshot’s platform and third-party providers including Together AI

On Moonshot’s own benchmark table, Kimi K2.5 Thinking scores 50.2 on HLE-Full with tools, ahead of GPT-5.2 xhigh (45.5) and Claude Opus 4.5 with extended thinking. That advantage inverts on HLE-Full without tools — GPT-5.2 leads at 34.5 vs Kimi’s 30.1 — which points to K2.5’s specific strength in agentic tool-use settings.

The Agent Swarm feature coordinates multiple K2.5 instances for long-horizon tasks. Moonshot’s docs describe it as a self-directed, swarm-style execution layer rather than a human-orchestrated multi-agent loop.

Grok 4.20 Beta: Pre-Release Evaluation

xAI’s Grok 4.20-beta-0309-reasoning has been added to the Document Arena under evaluation. The version label suggests a reasoning-mode build from early March 2026, running concurrently with Arena’s standard Grok 4 and Grok 4.1 entries. Musk set a June 2026 deadline for Grok to match Claude Opus 4.6; the Document Arena will provide one of the first third-party data points on whether that gap is closing.

Gemma 4 31B’s Document entry adds an important dimension: the model’s Arena ELO on text sits at 1452, placing it in the upper tier of open-weight models. Document performance will test whether Gemma 4’s general quality holds up on structured, multi-page inputs.

Image-to-WebDev: A New Benchmark Category

The Image-to-WebDev leaderboard is live but early. Arena’s methodology — blind human preference votes, with the model name hidden until a verdict is cast — provides the most contamination-resistant comparative ranking available. For front-end engineers, this will become a practical decision tool alongside Arena’s existing code and web leaderboards.

Results for all four additions are pending sufficient vote accumulation. Arena typically requires 500–1,000 pairwise comparisons per model pair before publishing a stable ELO estimate.