GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Muse Spark 1.2's 2.9-Point Terminal-Bench Gain Is a Harness Delta, Not a Model Delta

Meta released Muse Code and Muse Spark 1.2 on August 5, 2026. Coverage settled on one number: 82.9% on Terminal-Bench 2.1, framed as a meaningful improvement over the 80.0% Muse Spark 1.1 reported at its own launch in July. The framing is wrong, and Meta’s own methodology document explains why.

The Harness Shifted

Meta’s evaluation methodology specifies exactly which harness each model used:

For Terminal-Bench 2.1 and DeepSWE 1.1, each model is evaluated with its selected agent product: Muse Code for Muse Spark 1.2, mini-swe-agent for Muse Spark 1.1, Grok Build for Grok, Claude Code for Opus, Codex for GPT, Antigravity for Gemini, and Kimi Code for Kimi.

Muse Spark 1.1 ran in mini-swe-agent. Muse Spark 1.2 ran in Muse Code — a harness it was trained inside. Meta states this explicitly:

We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, alongside the integration of the Muse Code toolset to maximize harness compatibility.

The 2.9-point gain between the two generations is therefore a delta across a harness change and a training objective change simultaneously. It cannot be decomposed into model improvement versus harness advantage.

The Delta Is Smaller Than the Noise

The verified Terminal-Bench leaderboard contains pairs of identical models run in different harnesses. Filtering for rows where model, reasoning effort, benchmark version, and evaluator are all identical and only the harness differs yields four pairs:

ModelEffortVendor harnessTerminus 2Delta
GPT-5.5xhighCodex — 83.1%78.0%+5.1
Gemini 3 ProhighGemini CLI — 65.8%73.9%−8.1
Opus 4.7maxClaude Code — 68.9%66.1%+2.8
Gemini 3.1 ProhighGemini CLI — 65.8%65.6%+0.2

The signed deltas span −8.1 to +5.1. Mean absolute harness delta: 4.05 points.

Muse Spark 1.2’s reported generational gain is 2.9 points — smaller than the mean absolute harness delta, and smaller than three of the four individual pairs. The gain cannot be distinguished from harness noise using any data currently available.

No Independent Verification

The verified Terminal-Bench board as of early August showed 17 entries. Muse Spark 1.2 is not among them. Neither are Claude Opus 5, GPT-5.6 Sol, Kimi K3, or Gemini 3.6 Flash — five of the six models cited in Meta’s own comparison chart. The highest independently verified score on the board is 83.8% by Claude Fable 5 in Claude Code.

Meta’s 82.9% is a self-reported number from a self-selected harness, compared against a prior-generation measurement from a different harness. It is a claim about Muse Code as a system, not about Muse Spark 1.2 as a model.

Why It Matters Now

This is not a Meta-specific problem. The benchmark ecology for agent systems is broken at the seams. When a model is co-trained with a harness and then measured in that harness, the score travels with the system, not the model. The same dynamic will apply to any lab that tight-couples training and evaluation environments — and tight-coupling is now standard practice.

Terminal-Bench’s own verified entry for Muse Spark 1.1, run independently at xhigh effort in mini-swe-agent, landed at 76.2% ± 1.2% on July 9, 2026 — 3.8 points below Meta’s vendor figure, with the verified upper bound not reaching the vendor claim. That is the independent data point for the prior generation. There is no equivalent for the current one.

The GDPval-AA v2 and MCP Atlas figures Meta cited are not comparable either: those benchmarks use the provider’s own agent harnesses and are not CLI agent comparisons.

What the Numbers Actually Show

Stripping out the benchmark interpretation, Muse Spark 1.2 does show genuine progress on neutral third-party metrics: AA Intelligence Index rose 3 points to 54 (tied with GPT-5.5 at 55, Grok 4.5 at 54), and GDPval-AA v2 ELO rose 260 points to 1631, placing it fifth globally behind Claude Opus 5 (1852), GPT-5.6 Sol (1730), Kimi K3 (1685), and Claude Fable 5. Pricing held at $1.25/$4.25 per million input/output tokens.

The issue is specifically the Terminal-Bench gain, which received the most prominent headline treatment and is the one number that cannot be trusted without an independent controlled run. Until Muse Spark 1.2 appears on the verified leaderboard with a harness-independent evaluation, the 82.9% figure is a system score, not a model score.

Muse Spark 1.2 was added to Arena’s Agent Arena leaderboard on August 20, where head-to-head voting is harness-agnostic. That data will develop over time and provide the cleaner signal the vendor benchmarks cannot.