Arena Opens Image-to-WebDev Leaderboard; Kimi K2.5 and Grok 4.20 Beta Enter Document Rankings
Chatbot Arena added a new evaluation surface on April 15: the Image-to-WebDev leaderboard, which judges models on their ability to convert images and design mockups into working web applications. The category puts visual-to-code capability — a task that matters directly for front-end development workflows — under the same human-preference methodology Arena applies to its text and code rankings.
The same week, four models joined Arena’s Document leaderboard: Moonshot AI’s Kimi K2.5 Thinking, xAI’s Grok 4.20 Beta (reasoning), Google’s Gemma 4 31B, and Anthropic’s Claude Opus 4.6 Thinking. The Document arena tests models on multi-page document comprehension, extraction, and synthesis — tasks closer to real legal, financial, and research workflows than standard chat benchmarks.
Kimi K2.5 Thinking: 1 Trillion Parameters, Open-Sourced
Kimi K2.5 is Moonshot AI’s most capable open model. Released January 27, it is a Mixture-of-Experts architecture with 1 trillion total parameters and 32 billion active per forward pass — the same efficiency-focused MoE design used by Mixtral and Grok. Moonshot pretrained it on approximately 15 trillion mixed visual and text tokens, making multimodality native rather than bolt-on.
Key specs:
- Context window: 256K tokens
- Modes: Instant, Thinking (chain-of-thought), Agent, and Agent Swarm
- Architecture: 61-layer MoE, 400M-parameter MoonViT vision encoder
- Availability: Open weights on HuggingFace; API on Moonshot’s platform and third-party providers including Together AI
On Moonshot’s own benchmark table, Kimi K2.5 Thinking scores 50.2 on HLE-Full with tools, ahead of GPT-5.2 xhigh (45.5) and Claude Opus 4.5 with extended thinking. That advantage inverts on HLE-Full without tools — GPT-5.2 leads at 34.5 vs Kimi’s 30.1 — which points to K2.5’s specific strength in agentic tool-use settings.
The Agent Swarm feature coordinates multiple K2.5 instances for long-horizon tasks. Moonshot’s docs describe it as a self-directed, swarm-style execution layer rather than a human-orchestrated multi-agent loop.
Grok 4.20 Beta: Pre-Release Evaluation
xAI’s Grok 4.20-beta-0309-reasoning has been added to the Document Arena under evaluation. The version label suggests a reasoning-mode build from early March 2026, running concurrently with Arena’s standard Grok 4 and Grok 4.1 entries. Musk set a June 2026 deadline for Grok to match Claude Opus 4.6; the Document Arena will provide one of the first third-party data points on whether that gap is closing.
Gemma 4 31B’s Document entry adds an important dimension: the model’s Arena ELO on text sits at 1452, placing it in the upper tier of open-weight models. Document performance will test whether Gemma 4’s general quality holds up on structured, multi-page inputs.
Image-to-WebDev: A New Benchmark Category
The Image-to-WebDev leaderboard is live but early. Arena’s methodology — blind human preference votes, with the model name hidden until a verdict is cast — provides the most contamination-resistant comparative ranking available. For front-end engineers, this will become a practical decision tool alongside Arena’s existing code and web leaderboards.
Results for all four additions are pending sufficient vote accumulation. Arena typically requires 500–1,000 pairwise comparisons per model pair before publishing a stable ELO estimate.