GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Claude Sonnet 5 Thinking Enters Arena at 1551 Code ELO — Sonnet Tier Tops Every Flagship

Anthropic’s claude-sonnet-5-thinking entered Chatbot Arena on July 2, 2026, joining Code, Text, Search, Vision, and Document leaderboards simultaneously. Its opening code ELO of 1551 is the highest code score in the Stack Futures tracker — above every Opus-class and frontier model currently ranked.

The Numbers

ModelCode ELO
claude-sonnet-5-thinking1551
seed-2.1-pro-preview1539
kimi-k2.61514
claude-fable-51509
claude-opus-4.6 (max)1504
claude-opus-4.7 (max)1502
minimax-m31501
gpt-5.5-high1481
claude-opus-4.8 (max)1484

The 1551 code ELO puts claude-sonnet-5-thinking 42 points above Claude Fable 5 and 49 points above Claude Opus 4.7. Those two are Anthropic’s most capable publicly available models — Fable 5 launched at 80.3% SWE-Bench Pro, a benchmark no other model had cleared at launch. Sonnet 5 Thinking beats both on Arena’s crowd-sourced code preference votes.

Text, Search, Vision, and Document ELOs will populate as Arena accumulates head-to-head votes in those categories.

Why This Matters

The Sonnet tier has always sat below Opus in Anthropic’s model hierarchy: cheaper, faster, recommended for everyday coding and analysis rather than frontier-difficulty tasks. That positioning holds on raw capability benchmarks. It does not hold at Arena’s code ELO.

Extended thinking changes the calculation. Sonnet 5 with a reasoning budget can tackle problems that would previously have required Opus. The crowd-sourced votes at Arena are picking that up: users running real coding tasks are preferring the output.

The result is that model tier no longer maps cleanly to Arena ranking. A Sonnet-class model, with thinking enabled, is currently the top-ranked code model on the world’s largest AI evaluation platform.

Ticker Ceiling Note

Stack Futures’ ELO normalization clips at 1550. claude-sonnet-5-thinking at 1551 is the first model to exceed that ceiling since it was set. One model above the ceiling is not enough to trigger a ceiling review (the threshold is three or more). The clip is currently active for this model’s composite score calculation.

Composite Score Context

The composite score calculation partially offsets the ELO lead: pricing and agentic capability data for this model are not yet in the fundamentals pipeline, so both are imputed. That drops the composite to 790 (rank 4) despite the #1 code ELO. Once Anthropic publishes SWE-Bench Verified results and OpenRouter lists the model with pricing, the composite will update.

The ELO number itself is not imputed. It comes directly from Arena’s leaderboard feed.