GPT-6 Astra Tops Image-to-WebDev Arena at 1733 ELO, 23 Points Clear of Claude Fable 5.1
Arena’s Image-to-WebDev leaderboard added four models on September 13: GPT-6 Astra (Max), Claude Fable 5.1 (Max), Muse Spark 1.3 (Max), and GLM-5.3 Flash. GPT-6 Astra took the top spot, pushing Claude Fable 5.1 into second.
New Rankings
| Rank | Model | ELO |
|---|---|---|
| 1 | GPT-6 Astra (Max) | 1733 |
| 2 | Claude Fable 5.1 (Max) | 1710 |
| 3 | Claude Opus 5 (Max) | 1665 |
| 4 | Muse Spark 1.3 (Max) | 1645 |
| 6 | Claude Fable 5 | 1623 |
| 8 | GPT-5.6 Sol (xHigh) | 1604 |
| 10 | GLM-5.3 Flash | 1588 |
GPT-6 Astra at 1733 sits 23 points above Claude Fable 5.1 and 129 points above GPT-5.6 Sol (xHigh), its predecessor on the OpenAI lineup. Claude Fable 5.1 opens at 1710, 87 points above Fable 5 at 1623.
Meta’s Muse Spark 1.3 enters at 1645 ELO, placing fourth, while Z.ai’s GLM-5.3 Flash lands at 1588 in tenth place.
What Changed
The ELO ceiling on Image-to-WebDev has risen from 1627 in July to 1733 now, a 106-point expansion in two months as successive model generations have landed on the benchmark.
GPT-6 Astra’s top position here is independent of its showing on coding agent benchmarks, where it trails Fable 5.1 on LiveBench Agentic Coding (57.3 vs. 66.1). Image-to-WebDev measures models generating websites from images and screenshots alongside agentic coding workflows involving multi-step reasoning and tool use — a distinct capability surface from pure software engineering benchmarks.
The 129-point cross-generation jump from GPT-5.6 Sol to GPT-6 Astra on this leaderboard is notably large. GPT-5.6 Sol at 1604 was already a competitive model on this benchmark; GPT-6 Astra’s 1733 represents a step-change in web generation quality as judged by human evaluators.