Fable 5 Tops Arena's New Image-to-WebDev Leaderboard at 1627 — Anthropic Occupies All Top 7
Arena added claude-fable-5 and claude-sonnet-5-high to its Image-to-WebDev leaderboard on July 17. Both entered the top-10 immediately, extending Anthropic’s grip on the category: the top 7 positions are held by Claude models across four generations. GPT-5.5-xhigh is the first non-Anthropic entry, at #8 with 1525 ELO.
The leaderboard
| Rank | Model | ELO | Votes | Price |
|---|---|---|---|---|
| 1 | claude-fable-5 | 1,627 | 1,930 | $10/$50 per M |
| 2 | claude-opus-4-7-thinking | 1,581 | 4,282 | $5/$25 per M |
| 3 | claude-opus-4-7 | 1,567 | 4,618 | $5/$25 per M |
| 4 | claude-opus-4-6-thinking | 1,547 | 5,279 | $5/$25 per M |
| 5 | claude-sonnet-4-6 | 1,544 | 5,538 | $3/$15 per M |
| 6 | claude-opus-4-6 | 1,537 | 5,273 | $5/$25 per M |
| 7 | claude-sonnet-5-high | 1,533 | 1,492 | $2/$10 per M |
| 8 | gpt-5.5-xhigh | 1,525 | 4,210 | $5/$30 per M |
Fable 5’s 1,627 ELO is 46 points above Opus 4.7 thinking and 102 points above the first non-Anthropic entry. That is a more concentrated top-of-table than Anthropic holds on the text and code Arena leaderboards, where GPT-5.6 Sol, Gemini, and Grok models appear higher in the rankings.
Image-to-WebDev versus text-prompt frontend code
The distinction matters. Image-to-WebDev takes a screenshot or design image and asks the model to reproduce it as functional code. Text-prompt frontend code takes a natural-language description and asks the model to build the interface from scratch. They test different capabilities.
On text-prompt frontend code, Kimi K3 sits at #1 with 1,679 ELO — the result that prompted discussion about open Chinese models reaching the frontier. On image-to-webdev, Kimi K3 does not appear in the current top 8. The two leaderboards are ranking different skills.
Claude models have historically shown strong performance on tasks that combine visual analysis with structured code generation. Image-to-WebDev stacks those capabilities in sequence: understand the design, reproduce it precisely, produce valid HTML/CSS/JavaScript. Models that can hold a high-fidelity mental model of a reference image while generating code have an advantage here that does not transfer directly to the text-prompt category.
The Sonnet 5 High cost argument
Claude Sonnet 5 High at #7 with 1,533 ELO is the most cost-efficient model in the top tier. At $2/$10 per million input/output tokens, it costs one-fifth of Fable 5 ($10/$50) and less than half of Opus 4.7 ($5/$25). The 94-ELO gap between Sonnet 5 High and Fable 5 on this specific benchmark is the number that determines whether the price difference is worth it for any given workload.
GPT-5.5-xhigh at #8 (1,525 ELO, $5/$30) costs more than Sonnet 5 High and scores 8 ELO points below it, making it the weaker value proposition in this specific category at current prices.
Vote counts and rating stability
Fable 5 and Sonnet 5 High entered on July 17 with vote counts in the 1,400-1,900 range after 48 hours. The older models — Opus 4.6 has 5,000-plus votes — have more stable ratings. Fable 5’s 1,627 carries a margin of roughly ±15 ELO. The rank ordering at the top of the table has enough statistical separation to be meaningful, but the margins between ranks 3 through 7 (1,567 to 1,533) are narrow enough that additional votes could reorder them.
The frontier shift
Anthropic launched Fable 5 on June 9 and it is now accumulating Arena position across every category it has entered. Its text Arena ELO is competitive with the top closed models. Its code Arena position leads or follows closely behind GPT-5.6 Sol. Image-to-WebDev is a third leaderboard where it opens at #1.
The pattern across categories is consistent: Fable 5 has not been outranked by any model on any Arena leaderboard it has entered during its first few days of voting. How long that holds as vote counts stabilize — and as Kimi K3 accumulates evaluation history — will be the benchmark story to watch through July.