GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Claude Opus 4.8 Hits 1512 on Chatbot Arena — First Model to Break the 1510 Barrier

Claude Opus 4.8 was added to the Chatbot Arena (arena.ai) text, vision, document, and code leaderboards on June 6. It settled at 1512 ELO on the main text board — the first model to breach the 1510 tier on Chatbot Arena’s main text leaderboard.

Arena Text ELO: Current Top Tier

RankModelELO
1Claude Opus 4.81512
2GPT-5.5 (Pro)~1506
3Gemini 3.1 Pro Preview~1505
4Claude Opus 4.7~1505

The second through fourth ranked models are approximately interchangeable at 1505-1506. Opus 4.8 holds a 6-7 point lead that is narrow but clear.

Coding ELO: Wider Separation

The coding leaderboard shows more separation. Opus 4.8 leads at approximately 1582 Elo, ahead of Opus 4.7 at 1567 — a 15-point gap in the domain where Anthropic’s coding investment is most concentrated. FrontierSWE, the independent reproducibility tracker for SWE-bench results, ranks Opus 4.8 first at 2.74 ranking score versus GPT-5.5 at 3.06 and Opus 4.7 at 4.15, with the main improvement over 4.7 attributed to better consistency rather than peak capability.

What Changed From Opus 4.7

Opus 4.7 entered Arena in April and settled at 1505 ELO within three days. Opus 4.8’s 1512 debut represents a 7-point gain over that baseline. On Artificial Analysis’s GDPval-AA Elo — a different leaderboard measuring real-world knowledge and productivity delivery — Opus 4.8 sits at 1890, implying a 66.7% pairwise win rate against GPT-5.5.

The AA Intelligence Index (separate from Arena ELO) has Opus 4.8 at 61, leading GPT-5.5 at 60 and Gemini 3.1 Pro at 57.

The Caveat

Toolathon — real-world tool-use task evaluations — shows marginal improvement: Opus 4.8 at 59.9% versus Opus 4.7’s previous high of 59.3%. For agentic work requiring sustained tool use, the upgrade is small. The coding and general intelligence gains are real; the tool-use gains are not yet.

The one benchmark where GPT-5.5 still leads: Terminal-Bench 2.1 at 78.2% versus Opus 4.8 at 74.6%. Terminal-heavy agent workflows should still factor that in before switching.

Pattern Context

Anthropic has now landed three successive models at or above 1505 Arena ELO within days of launch — Opus 4.6 (thinking variant, 1504), Opus 4.7 (1505), Opus 4.8 (1512). Each adds 5-7 points. At current trajectory, the next Anthropic flagship enters somewhere around 1517-1520. The Arena ceiling has not plateaued.