GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

GPT-6 Astra Tops Image-to-WebDev Arena at 1733 ELO, 23 Points Clear of Claude Fable 5.1

Arena’s Image-to-WebDev leaderboard added four models on September 13: GPT-6 Astra (Max), Claude Fable 5.1 (Max), Muse Spark 1.3 (Max), and GLM-5.3 Flash. GPT-6 Astra took the top spot, pushing Claude Fable 5.1 into second.

New Rankings

RankModelELO
1GPT-6 Astra (Max)1733
2Claude Fable 5.1 (Max)1710
3Claude Opus 5 (Max)1665
4Muse Spark 1.3 (Max)1645
6Claude Fable 51623
8GPT-5.6 Sol (xHigh)1604
10GLM-5.3 Flash1588

GPT-6 Astra at 1733 sits 23 points above Claude Fable 5.1 and 129 points above GPT-5.6 Sol (xHigh), its predecessor on the OpenAI lineup. Claude Fable 5.1 opens at 1710, 87 points above Fable 5 at 1623.

Meta’s Muse Spark 1.3 enters at 1645 ELO, placing fourth, while Z.ai’s GLM-5.3 Flash lands at 1588 in tenth place.

What Changed

The ELO ceiling on Image-to-WebDev has risen from 1627 in July to 1733 now, a 106-point expansion in two months as successive model generations have landed on the benchmark.

GPT-6 Astra’s top position here is independent of its showing on coding agent benchmarks, where it trails Fable 5.1 on LiveBench Agentic Coding (57.3 vs. 66.1). Image-to-WebDev measures models generating websites from images and screenshots alongside agentic coding workflows involving multi-step reasoning and tool use — a distinct capability surface from pure software engineering benchmarks.

The 129-point cross-generation jump from GPT-5.6 Sol to GPT-6 Astra on this leaderboard is notably large. GPT-5.6 Sol at 1604 was already a competitive model on this benchmark; GPT-6 Astra’s 1733 represents a step-change in web generation quality as judged by human evaluators.