Muse Spark 1.2's 2.9-Point Terminal-Bench Gain Is a Harness Delta, Not a Model Delta
Meta released Muse Code and Muse Spark 1.2 on August 5, 2026. Coverage settled on one number: 82.9% on Terminal-Bench 2.1, framed as a meaningful improvement over the 80.0% Muse Spark 1.1 reported at its own launch in July. The framing is wrong, and Meta’s own methodology document explains why.
The Harness Shifted
Meta’s evaluation methodology specifies exactly which harness each model used:
For Terminal-Bench 2.1 and DeepSWE 1.1, each model is evaluated with its selected agent product: Muse Code for Muse Spark 1.2, mini-swe-agent for Muse Spark 1.1, Grok Build for Grok, Claude Code for Opus, Codex for GPT, Antigravity for Gemini, and Kimi Code for Kimi.
Muse Spark 1.1 ran in mini-swe-agent. Muse Spark 1.2 ran in Muse Code — a harness it was trained inside. Meta states this explicitly:
We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, and subagents, alongside the integration of the Muse Code toolset to maximize harness compatibility.
The 2.9-point gain between the two generations is therefore a delta across a harness change and a training objective change simultaneously. It cannot be decomposed into model improvement versus harness advantage.
The Delta Is Smaller Than the Noise
The verified Terminal-Bench leaderboard contains pairs of identical models run in different harnesses. Filtering for rows where model, reasoning effort, benchmark version, and evaluator are all identical and only the harness differs yields four pairs:
| Model | Effort | Vendor harness | Terminus 2 | Delta |
|---|---|---|---|---|
| GPT-5.5 | xhigh | Codex — 83.1% | 78.0% | +5.1 |
| Gemini 3 Pro | high | Gemini CLI — 65.8% | 73.9% | −8.1 |
| Opus 4.7 | max | Claude Code — 68.9% | 66.1% | +2.8 |
| Gemini 3.1 Pro | high | Gemini CLI — 65.8% | 65.6% | +0.2 |
The signed deltas span −8.1 to +5.1. Mean absolute harness delta: 4.05 points.
Muse Spark 1.2’s reported generational gain is 2.9 points — smaller than the mean absolute harness delta, and smaller than three of the four individual pairs. The gain cannot be distinguished from harness noise using any data currently available.
No Independent Verification
The verified Terminal-Bench board as of early August showed 17 entries. Muse Spark 1.2 is not among them. Neither are Claude Opus 5, GPT-5.6 Sol, Kimi K3, or Gemini 3.6 Flash — five of the six models cited in Meta’s own comparison chart. The highest independently verified score on the board is 83.8% by Claude Fable 5 in Claude Code.
Meta’s 82.9% is a self-reported number from a self-selected harness, compared against a prior-generation measurement from a different harness. It is a claim about Muse Code as a system, not about Muse Spark 1.2 as a model.
Why It Matters Now
This is not a Meta-specific problem. The benchmark ecology for agent systems is broken at the seams. When a model is co-trained with a harness and then measured in that harness, the score travels with the system, not the model. The same dynamic will apply to any lab that tight-couples training and evaluation environments — and tight-coupling is now standard practice.
Terminal-Bench’s own verified entry for Muse Spark 1.1, run independently at xhigh effort in mini-swe-agent, landed at 76.2% ± 1.2% on July 9, 2026 — 3.8 points below Meta’s vendor figure, with the verified upper bound not reaching the vendor claim. That is the independent data point for the prior generation. There is no equivalent for the current one.
The GDPval-AA v2 and MCP Atlas figures Meta cited are not comparable either: those benchmarks use the provider’s own agent harnesses and are not CLI agent comparisons.
What the Numbers Actually Show
Stripping out the benchmark interpretation, Muse Spark 1.2 does show genuine progress on neutral third-party metrics: AA Intelligence Index rose 3 points to 54 (tied with GPT-5.5 at 55, Grok 4.5 at 54), and GDPval-AA v2 ELO rose 260 points to 1631, placing it fifth globally behind Claude Opus 5 (1852), GPT-5.6 Sol (1730), Kimi K3 (1685), and Claude Fable 5. Pricing held at $1.25/$4.25 per million input/output tokens.
The issue is specifically the Terminal-Bench gain, which received the most prominent headline treatment and is the one number that cannot be trusted without an independent controlled run. Until Muse Spark 1.2 appears on the verified leaderboard with a harness-independent evaluation, the 82.9% figure is a system score, not a model score.
Muse Spark 1.2 was added to Arena’s Agent Arena leaderboard on August 20, where head-to-head voting is harness-agnostic. That data will develop over time and provide the cleaner signal the vendor benchmarks cannot.