GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

LiveBench September 2026: Fable 5.1 Leads at 83.4, Muse Spark 1.3 Scores 81.6 at $0.22 per Task

September’s LiveBench leaderboard adds three new entrants to a benchmark that has been reshuffled repeatedly since mid-year. Claude Fable 5.1 holds the top position at 83.4. Meta’s Muse Spark 1.3 is the efficiency story.

Current Standings

RankModelOverallAgentic CodingCost/Task
1Claude Fable 5.1 (max)83.466.1$1.21
2Claude Fable 5 (max)83.062.2$1.44
3Muse Spark 1.3 (xhigh)81.664.1$0.22
4GPT-5.6 Sol (max)81.056.2$0.52
5GPT-5.5 Thinking (xhigh)80.254.0$0.44
6Claude Opus 5 Thinking (max)80.165.2—

Fable 5.1’s Edge

Fable 5.1 improves over Fable 5 by 0.4 points overall, but the substantive gain is in agentic coding: 66.1 versus 62.2. On reasoning (91.7) and mathematics (97.0), Fable 5.1 leads all measured models. It costs $1.21 per successful task, compared to $1.44 for Fable 5 — a 16% cost reduction for the successor model, which also outperforms the original.

Claude Opus 5 Thinking ranks sixth at 80.1 overall, but leads the agentic coding column at 65.2%, slightly above Fable 5.1’s 66.1%. The difference is within margin of measurement, but it positions Opus 5 as the better agentic coding model for those who can access it at extended compute levels.

Muse Spark 1.3: The Cost Outlier

Muse Spark 1.3 at extended effort (xhigh) posts an 81.6 overall score at $0.22 per successful task. Compared to the adjacent models:

  • $0.22 vs $0.52 for GPT-5.6 Sol at a higher overall score (81.6 vs 81.0)
  • $0.22 vs $1.21 for Claude Fable 5.1 at 1.8 points below Fable 5.1’s overall score
  • Agentic coding: 64.1% — 7.9 points above Sol (56.2%) at less than half the cost

For production workloads that do not require peak quality but do require cost predictability at scale, Muse Spark 1.3’s positioning is difficult to ignore. It is the only model in the top 6 that delivers over 80.0 overall under $0.30 per task.

The model also scores above every OpenAI entry in the reasoning column (89.7), tied with Claude Fable 5 and GPT-5.5 Thinking. Meta released Muse Spark 1.3 as an open-weight model; running it self-hosted eliminates the per-task cost entirely for organizations with available GPU capacity.

Agentic Coding Is the Dividing Line

The agentic coding column continues to be where the frontier separates most visibly. Models that cluster between 80-83 overall diverge sharply on agentic coding:

  • Fable 5.1: 66.1
  • Muse Spark 1.3: 64.1
  • Fable 5: 62.2
  • GPT-5.6 Sol: 56.2
  • GPT-5.5 Thinking: 54.0

The 10-point gap between the Anthropic Fable family and GPT-5.6 Sol on this single column, despite being within 2.4 points on overall score, suggests the benchmark is correctly isolating a real capability difference. LiveBench’s agentic coding tasks require multi-step code execution and debugging rather than single-shot completion.

GPT-6 Astra

GPT-6 Astra, which launched September 3, does not yet appear in the visible LiveBench top 10. Artificial Analysis’s Coding Agent Index benchmarks (published the same day) put Astra at 67 in the coding agent evaluation — above Sol’s equivalent — suggesting it will likely enter the LiveBench agentic coding column above 56.2 when results are added. Its position on the overall LiveBench ranking will depend on whether the token efficiency gains that benefit Astra in coding-specific harnesses translate to the broader evaluation suite.

LiveBench runs live evaluations to prevent training set contamination. Results update as new models are submitted; scores shown here are from the leaderboard as of September 4, 2026.