GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 827 -5.3%
QWEN-38X 824 —
CL-OP55X 820 —
GPT-6A 820 —
GROK-46H 820 -5.2%
GLM-5 784 -8.4%
KIMI-K3X 742 -8.4%
CL-FAB5H 742 -5.7%
CL-OP5H 718 -6%
CL-OP5X 708 -18.2%
CL-OP46H 696 -6.2%
CL-OP47H 688 -6.1%
GEM-38FH 677 +0.1%
GEM-37FH 655 -24.3%
GPT-56S 619 —
GPT-55H 580 —
CL-OP47 579 -0.7%
INKL 531 —
GEM-31P 512 —
GEM-3P 498 —
CL-OP46 496 —
CL-OP48 489 -0.2%
← Back to feed

Arena's Seven Web Dev Categories Reveal Which Model Actually Wins Your Use Case

The Code Arena leaderboard produces one number. Arena’s new category-level breakdown, published May 8, reveals how much that number conceals.

Breaking six months of web development votes (November 2025 to April 2026) into seven concrete categories, Arena found that model rankings invert depending on which type of web work is being evaluated. The methodology uses multi-label tagging — a single vote can carry multiple category tags — and the resulting data shows that task type matters as much as raw capability when selecting a model for production work.

The Seven Categories

Arena’s taxonomy covers the web development tasks users actually build:

  1. Brand marketing and informational websites
  2. Data and analytics applications
  3. Consumer product and platform applications
  4. Gaming
  5. Simulations
  6. Content creation and editing tools
  7. Reference-based design

The distribution has shifted over the six-month window. Practical, real-world categories — brand marketing, data/analytics, and consumer product apps — have grown as a share of total web dev prompts. Gaming and simulations have declined in relative share while remaining meaningful in absolute volume.

Proprietary Model Breakdown

Claude Opus 4.7 Thinking (Anthropic) shows consistently high rankings across all seven categories. In Arena’s radar plot, its profile shows no obvious weak spot — the broadest model for general web development. For teams that need one model to cover multiple use-case types without category-by-category selection, it is the default answer.

GPT-5.5 High (OpenAI) performs strongly overall but diverges most from Claude in interactive categories. Gaming and simulations are its visible standout — OpenAI’s architecture appears to translate especially well to the generation of interactive, stateful, and game-logic-heavy applications. The gap narrows in practical website and product categories.

Muse Spark (Meta) stands out among proprietary models in the categories that have grown fastest: brand marketing and informational websites, reference-based design, and consumer product and platform applications. Its profile is more concentrated than Claude’s — strong where practical website and product-building tasks dominate, less competitive in the interactive categories where GPT-5.5 leads.

Open-Source Picture

GLM-5.1 (Zhipu AI) and Kimi K2.6 (Moonshot AI) both achieve broad category coverage in the open-source rankings, with different areas of emphasis that Arena’s radar plots make legible for the first time. The difference between their profiles had been invisible in aggregate ELO.

Gemma-4-31B (Google) shows stronger performance in consumer product and platform applications than its aggregate ranking implies. Category-level analysis makes this competitive position visible.

Mimo-V2.5 Pro (Xiaomi) leads among open-source models in gaming, simulations, and content creation and editing — interactive and creative categories where its profile is most differentiated.

Why Aggregate Rankings Lie

Arena’s analysis makes explicit what practitioners already know: model selection is a function of task type, not just overall quality. A model ranked fourth on the global leaderboard may be ranked first for the specific category of work a team is doing. Without category-level data, that information is inaccessible.

The practical implication is a shift in how development teams should evaluate models. Aggregate Arena ELO is a reasonable starting point. Category-level rankings are the finishing step for any deployment decision where the use case is specific.

Arena plans further refinements: finer subcategories, tool-use pattern analysis, and breakdowns by interaction complexity, execution success, and visual quality. The current seven-category taxonomy is a first pass — and already enough to show that the one-number leaderboard leaves material information on the floor.