GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Arena July 27: Claude Opus 5 High Enters Text and Code Leaderboards, Kimi K3 and Inkling Crack Agent Arena

Arena’s July 27, 2026 leaderboard changelog is one of the denser single-day additions of the year. Five models across six leaderboards, split across two distinct competitive contexts: the standard text and vision boards where Claude Opus 5 High now competes, and the Agent Arena where Kimi K3 and Inkling make their first appearances.

Claude Opus 5 High on Text, Document, and Code Arena

Claude Opus 5 High — the extended-thinking, higher-effort variant of Anthropic’s Opus 5 family — enters Arena’s Text, Document, and Code leaderboards simultaneously. The Artificial Analysis Intelligence Index currently places Opus 5 (max) and Opus 5 (xhigh) as the two highest-intelligence models it tracks, above Fable 5 and GPT-5.6 Sol.

Arena ELO rankings for Opus 5 High will take time to calibrate as battle volume accumulates, but its benchmark position coming in is the strongest of any model added to Arena this year. On SWE-bench Verified, Opus 5 posts 96.0% — above Fable 5’s 95.0% and Mythos 5’s 95.5%. On ARC-AGI-3, it holds the current record at 30.2%, four times the next-best score from GPT-5.6 Sol at 7.8%. On the AA-Briefcase agentic knowledge work benchmark, it posts Elo 1720, 146 points clear of Fable 5.

What the Arena battles will test is whether that benchmark advantage translates into human preference ratings across open-ended tasks — the domain where Fable 5 has historically been strong and where Arena’s methodology measures what the benchmarks cannot.

Kimi K3 and Inkling Join Agent Arena

The more structurally significant addition may be the Agent Arena entries. Kimi K3 and Inkling both entered the Agent Arena on July 27 — the leaderboard Arena launched in June with causal tracing methodology to evaluate models on real multi-step tasks rather than single-turn responses.

Kimi K3 brings the strongest open-weight benchmark profile currently on record for an agent model: 93.4% SWE-bench Verified, #3 globally behind only Mythos 5 and Fable 5. On Harvey LAB-AA, a 120-task private legal benchmark, it posted 26.7% — nearly double Fable 5’s score, and it resolved 15 bugs Fable refused to touch. On Frontend Code Arena, it already holds the top position at 1679 Elo, the first Chinese model to beat both Fable 5 and GPT-5.6.

Inkling, from Thinking Machines Lab, enters with a 975B MoE architecture at 41B active parameters under Apache 2.0. Its SWE-bench Verified score of 77.6% places it in the upper-middle tier — above most proprietary models from six months ago, competitive with current Gemini Flash and GPT-5.4 variants. Its Agent Arena position is not expected to challenge the top of the table, but its open-weight status and permissive licensing make it a reference point for what self-hosted agent infrastructure can do.

Grok 4.5 and Muse Spark 1.1 on Vision and Document

Grok 4.5 joined the Vision and Document leaderboards. Already rated on Artificial Analysis’ Intelligence Index at 62 and priced at $2/$6 per million tokens, it enters vision evaluation with known agentic credentials — 86.6% SWE-bench Verified and the top WANDR score at 0.328, beating Opus 4.8.

Meta’s Muse Spark 1.1 enters Vision and Document, updated from the original Muse Spark that debuted at #4 on the Intelligence Index in April. The 1.1 designation and developer API launched alongside this Arena addition; specific benchmark improvements over the original have not been independently confirmed.

The Agent Arena Picture

After the July 27 additions, the Agent Arena now includes the full tier of frontier proprietary models plus the two most capable open-weight entries. The competitive map coming in: Fable 5 at the top (12.94% composite), followed by GPT-5.5 High variants, with Opus 4.8 Thinking at rank 2. Kimi K3’s entry sets up the first credible open-weight challenge to proprietary dominance of the agentic leaderboard.

How Kimi K3 performs relative to its SWE-bench Verified rank — which would put it near the top of any agentic capability ranking — versus its Agent Arena composite score will be the benchmark story to watch as battle volume builds.