GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Meta's Muse Spark 1.3 Hits #4 on LiveBench at $0.22 per Task

Muse Spark 1.3, from Meta, has entered the September 2026 LiveBench leaderboard at fourth place overall with a score of 81.6 — putting it ahead of every Google model in the general ranking, and within 1.8 points of the leader, Claude Fable 5.1.

Its cost-per-successful-task is $0.219, running xHigh Effort. That is less than a fifth of Claude Fable 5.1’s $1.212, and roughly a third of GPT-6 Astra’s $0.736.

Where It Sits in the Top 5

The September 2026 LiveBench top 5:

ModelOverallAgentic CodingCost/Task
Claude Fable 5.1 (max)83.466.1$1.212
Claude Fable 5 (max)83.062.2$1.439
GPT-6 Astra (max)82.257.3$0.736
Muse Spark 1.3 (xHigh)81.664.1$0.219
DeepSeek V4.1 Flash (max)81.177.3$0.029

Muse Spark 1.3’s 81.6 puts it 0.6 points behind GPT-6 Astra in the overall standings, but ahead of every Google model in the ranking.

Category Breakdown

The model scores 89.7 on Reasoning, matching Claude Fable 5. GPT-6 Astra leads on Reasoning at 92.7 — the one category where OpenAI holds the top position. In Mathematics it hits 95.9, in Coding 81.1, in Language 82.8, and in Instruction Following 78.0.

Its Agentic Coding score of 64.1 places it third among paid frontier models behind Fable 5.1 (66.1) and Claude Opus 5 Thinking (65.2) — well ahead of GPT-6 Astra’s 57.3 and Fable 5’s 62.2. Only DeepSeek V4.1 Flash (77.3) scores higher on that specific column, but Flash is open-weight and costs $0.029/task.

Cost-Performance Context

The case for Muse Spark 1.3 is largely economic. A customer running at scale who does not need Anthropic’s Agentic Coding edge (Fable 5.1 leads by 2 points on that column) can get comparable overall performance for roughly one-fifth the inference cost.

The gap between Muse Spark 1.3 (81.6, $0.219) and Claude Fable 5.1 (83.4, $1.212) is 1.8 LiveBench points and 5.5x in cost. Whether that tradeoff is worthwhile depends on the task mix, but at high volume the arithmetic is significant.

Lab Context

Meta appears in LiveBench’s organizational filter alongside Anthropic, OpenAI, Google, DeepSeek, and xAI. Muse Spark 1.3 was also added to the Arena.ai leaderboard in the September 8 update, per the Arena changelog, alongside GPT-6 Astra Max and the ChatGPT Images 2.5 generation.

The model has not been subject to independent SWE-bench Verified or tau-bench evaluation as of this writing. Its LiveBench scores are the primary external signal available. Agentic capability scoring on the Stack Futures ticker remains pending until primary benchmark data arrives.