GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

AA Intelligence Index July 2026: GPT-5.6 Sol Is 1 Point Behind Fable 5, Four Models from Three Labs Tied at 51

The benchmark leaderboard that most enterprise developers use for model selection has settled into two distinct tiers. At the top, Claude Fable 5 holds a 1-point lead over GPT-5.6 Sol Max on Artificial Analysis’s Intelligence Index — 60 to 59. That gap is within the index’s statistical uncertainty. Below them, a five-point drop separates Opus 4.8 at 56 and Grok 4.5 at 54 from a four-model cluster at 51 that spans three separate labs.

The Current Standings

ModelScoreLab
Claude Fable 5 (max, with fallback)60Anthropic
GPT-5.6 Sol (max)59OpenAI
Claude Opus 4.8 (max)56Anthropic
Grok 4.5 (high)54xAI
GLM-5.2 (max)51Z.ai
GPT-5.4 (xhigh)51OpenAI
GPT-5.6 Luna (max)51OpenAI
Muse Spark 1.1 (xhigh)51Meta

Fable 5’s Compressed Lead

When Fable 5 launched in June 2026, it posted 64.9 on the AA Intelligence Index — roughly 6 points clear of the field. Artificial Analysis subsequently overhauled the index, replacing τ²-Bench Telecom with τ³-Bench Banking as a component benchmark. The methodology change shifted absolute scores across the board. Post-overhaul, Fable 5 scores 60. GPT-5.6 Sol, evaluated on the new methodology from day one, scores 59.

AA’s own note on the current leaderboard states that “Claude Fable 5 (with fallback) and GPT-5.6 Sol (max) are the highest intelligence models, followed by GPT-5.6 Sol (xhigh) and GPT-5.6 Sol (high).” The three GPT-5.6 Sol variants occupy second, third, and fourth in the column before Opus 4.8 enters at 56.

What Changed at 57

A year ago, the defining AA leaderboard story was a three-way tie at 57: Claude Opus 4.7, Gemini 3.1 Pro Preview, and GPT-5.4 were indistinguishable on the index within its confidence interval. The practical implication at the time was that model selection had to shift from headline composite scores to workload-specific strengths.

Three model generations later, the models that headlined the “tie at 57” story now anchor the cluster at 51. GPT-5.4 (xhigh) still scores 51 — held in place as the frontier moved up by 9 points and left it behind. The ceiling rose by roughly 3 points per quarter.

The Cluster at 51

Four models from three different labs score 51 simultaneously. That number includes:

  • GLM-5.2 (max) — Z.ai’s Chinese open-weights model, which leads open-weights rankings and has separately topped Code Arena web dev categories
  • GPT-5.4 (xhigh) — OpenAI’s previous generation, still competitive for teams with established GPT-5.4 integrations
  • GPT-5.6 Luna (max) — OpenAI’s mid-tier Sol variant, positioned between Sol and the base
  • Muse Spark 1.1 (xhigh) — Meta’s first developer-API model, launched July 9 at $1.25/$4.25 per million tokens

The practical difference between these four models is cost, not intelligence index. Artificial Analysis evaluated Muse Spark 1.1 on the full index suite and found it consumed 94M output tokens — less than GPT-5.4 (109M), GPT-5.6 Luna (125M), or GLM-5.2 (141M) to reach the same score. That translates to approximately $0.26 per Intelligence Index task for Muse Spark 1.1 against materially higher costs for its three tied peers.

On Humanity’s Last Exam, Muse Spark 1.1 scores 45% — within 1 point of Claude Opus 4.8 (46%) and ahead of GPT-5.5 (44%) and Grok 4.5 (40%) on the specific test. The overall 51 on the AA index understates its strength on the hardest knowledge-retrieval tasks.

The Top-Tier Decision

The 1-point gap between Fable 5 and GPT-5.6 Sol at the 60/59 level is not actionable as a pure benchmark selection criterion. On LiveBench, the models differentiate sharply by workload: Fable 5 leads on language (90.7%) and instruction following (75.8%), while GPT-5.6 Sol leads overall and both GPT-5.6 variants lead on agentic coding. Cost separates them further — Fable 5 Max Effort runs $1.573 per successful LiveBench task against GPT-5.6 Sol Max at $0.589.

The Intelligence Index at 60 vs 59 is a tie for practical purposes. The workload-level benchmarks are where the actual selection decision happens.