LiveBench September 2026: Fable 5.1 Leads at 83.4, Muse Spark 1.3 Scores 81.6 at $0.22 per Task
September’s LiveBench leaderboard adds three new entrants to a benchmark that has been reshuffled repeatedly since mid-year. Claude Fable 5.1 holds the top position at 83.4. Meta’s Muse Spark 1.3 is the efficiency story.
Current Standings
| Rank | Model | Overall | Agentic Coding | Cost/Task |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 (max) | 83.4 | 66.1 | $1.21 |
| 2 | Claude Fable 5 (max) | 83.0 | 62.2 | $1.44 |
| 3 | Muse Spark 1.3 (xhigh) | 81.6 | 64.1 | $0.22 |
| 4 | GPT-5.6 Sol (max) | 81.0 | 56.2 | $0.52 |
| 5 | GPT-5.5 Thinking (xhigh) | 80.2 | 54.0 | $0.44 |
| 6 | Claude Opus 5 Thinking (max) | 80.1 | 65.2 | — |
Fable 5.1’s Edge
Fable 5.1 improves over Fable 5 by 0.4 points overall, but the substantive gain is in agentic coding: 66.1 versus 62.2. On reasoning (91.7) and mathematics (97.0), Fable 5.1 leads all measured models. It costs $1.21 per successful task, compared to $1.44 for Fable 5 — a 16% cost reduction for the successor model, which also outperforms the original.
Claude Opus 5 Thinking ranks sixth at 80.1 overall, but leads the agentic coding column at 65.2%, slightly above Fable 5.1’s 66.1%. The difference is within margin of measurement, but it positions Opus 5 as the better agentic coding model for those who can access it at extended compute levels.
Muse Spark 1.3: The Cost Outlier
Muse Spark 1.3 at extended effort (xhigh) posts an 81.6 overall score at $0.22 per successful task. Compared to the adjacent models:
- $0.22 vs $0.52 for GPT-5.6 Sol at a higher overall score (81.6 vs 81.0)
- $0.22 vs $1.21 for Claude Fable 5.1 at 1.8 points below Fable 5.1’s overall score
- Agentic coding: 64.1% — 7.9 points above Sol (56.2%) at less than half the cost
For production workloads that do not require peak quality but do require cost predictability at scale, Muse Spark 1.3’s positioning is difficult to ignore. It is the only model in the top 6 that delivers over 80.0 overall under $0.30 per task.
The model also scores above every OpenAI entry in the reasoning column (89.7), tied with Claude Fable 5 and GPT-5.5 Thinking. Meta released Muse Spark 1.3 as an open-weight model; running it self-hosted eliminates the per-task cost entirely for organizations with available GPU capacity.
Agentic Coding Is the Dividing Line
The agentic coding column continues to be where the frontier separates most visibly. Models that cluster between 80-83 overall diverge sharply on agentic coding:
- Fable 5.1: 66.1
- Muse Spark 1.3: 64.1
- Fable 5: 62.2
- GPT-5.6 Sol: 56.2
- GPT-5.5 Thinking: 54.0
The 10-point gap between the Anthropic Fable family and GPT-5.6 Sol on this single column, despite being within 2.4 points on overall score, suggests the benchmark is correctly isolating a real capability difference. LiveBench’s agentic coding tasks require multi-step code execution and debugging rather than single-shot completion.
GPT-6 Astra
GPT-6 Astra, which launched September 3, does not yet appear in the visible LiveBench top 10. Artificial Analysis’s Coding Agent Index benchmarks (published the same day) put Astra at 67 in the coding agent evaluation — above Sol’s equivalent — suggesting it will likely enter the LiveBench agentic coding column above 56.2 when results are added. Its position on the overall LiveBench ranking will depend on whether the token efficiency gains that benefit Astra in coding-specific harnesses translate to the broader evaluation suite.
LiveBench runs live evaluations to prevent training set contamination. Results update as new models are submitted; scores shown here are from the leaderboard as of September 4, 2026.