GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.6 Sol Takes Design Arena #1 at Elo 1353, 60 Points Clear of GPT-5.5

GPT-5.6 Sol has reached Elo 1353 on Design Arena, taking the top position ahead of Claude Fable 5 and GLM-5.2. The jump represents 18 leaderboard positions and 60 Elo points over GPT-5.5, making it the largest single-generation gain OpenAI has posted on this leaderboard.

What Design Arena Measures

Design Arena is distinct from code arenas and agent benchmarks. Models receive a prompt to create a webpage and produce the result without an agentic tool loop — no iterations, no browser feedback, one shot. Human voters compare anonymized outputs in pairwise matchups and pairwise preferences are converted to Elo ratings through standard statistical modeling. A higher score reflects visual preference, not software correctness.

That distinction matters. Models that win on SWE-bench or agent evals do not automatically win on design. Fable 5 holds Elo 1298 here — below GPT-5.6 Sol’s 1353 and in the same performance band as GLM-5.2. Anthropic’s model leads the AA Intelligence Index at 64.9 and SWE-bench Pro at 80.3%. On frontend visual quality as judged by human preference, GPT-5.6 Sol is ahead.

Why GPT-5.6 Sol Wins on Design

The evaluators attribute GPT-5.6 Sol’s gain to its rendered-page inspection capability. The model can examine what a page actually looks like after rendering — not just the markup — which lets it correct visual problems before submitting a final output. Rivals at comparable Elo scores require more interaction turns to reach similar preference rates.

That inspection advantage also feeds into speed. GPT-5.6 Sol establishes the new Pareto frontier for preference versus response time: it is faster than every other model at the 1353 Elo tier. The combination — higher preference scores and shorter response times — is what pushed it 18 positions up the rankings from GPT-5.5’s prior position.

Competitive Context

The Design Arena rankings as of July 13, 2026:

RankModelElo
1GPT-5.6 Sol1353
2GLM-5.2~1301
3Claude Fable 5~1298
4GPT-5.5~1293

GLM-5.2 and Fable 5 remain in the same performance band. GPT-5.6 Sol’s 55-point gap to the next cluster is the widest margin at the top of this leaderboard.

What It Does Not Change

This result is leaderboard-specific. On AA Intelligence Index, Fable 5 leads at 64.9. On SWE-bench Pro, Fable 5 leads at 80.3%. On Agent Arena, Fable 5 leads. The Design Arena result reflects a single capability: producing webpages humans prefer, once, from a prompt. For development workflows, agentic coding, and multi-step reasoning, the relative standings look different.

For teams building frontend-heavy tools or prototyping UI at scale, GPT-5.6 Sol has now demonstrated a measurable edge in visual output quality over every rival.