GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.6 Terra Leads Every Top-Tier Model on LiveBench Agentic Coding — Including GPT-5.6 Sol

The LiveBench leaderboard hides a counterintuitive result. GPT-5.6 Terra Max Effort — positioned as the middle tier in OpenAI’s three-model lineup — scores higher than GPT-5.6 Sol on agentic coding, the one LiveBench category most directly tied to agent task completion. That holds across the full top-five ranking: no other frontier model beats Terra’s 68.0% on that dimension.

The Numbers

ModelOverallAgentic CodingCost/Task
GPT-5.6 Sol Max Effort82.465.6%$0.589
Claude Fable 5 Max Effort80.846.9%$1.573
GPT-5.5 Thinking xHigh79.952.1%$0.530
GPT-5.6 Terra Max Effort79.868.0%$0.497
Claude Opus 4.8 Thinking xHigh78.956.1%$0.688
GPT-5.4 Thinking xHigh78.053.8%$0.387

Terra ranks fourth overall. It ranks first on agentic coding by 2.4 points over Sol — the model that sits 2.6 points above it in composite score — and by nearly 12 points over Claude Fable 5.

Where Sol Wins

Sol’s composite advantage is real and comes from every other category. Sol leads Terra on:

  • Reasoning: 91.7% vs 90.6% (+1.1pp)
  • Coding (non-agentic): 83.9% vs 78.2% (+5.7pp)
  • Mathematics: 96.2% vs 94.9% (+1.3pp)
  • Language: 87.7% vs 82.9% (+4.8pp)
  • Instruction following: 71.8% vs 64.6% (+7.2pp)

The 2.6-point overall gap between Sol (82.4) and Terra (79.8) is not noise. For general-purpose tasks — writing, reasoning, mathematics, structured instruction execution — Sol is the stronger model.

Why the Agentic Inversion Matters

Agentic coding in LiveBench measures sustained, multi-step software execution: understanding a codebase, decomposing a task, running multiple tool calls, verifying output, and recovering from failures. It does not measure single-turn coding skill, which is where Sol’s 5.7-point lead in the standard coding category comes from.

The divergence between Sol and Terra on those two coding categories — Terra wins agents (+2.4pp), Sol wins code (+5.7pp) — suggests their training distributions differ in how they handle sequential execution vs. direct generation. Sol appears to optimise for the latter more aggressively, at the cost of something in the orchestration layer.

Terra also beats Sol on data analysis (79.3% vs 79.8%)… by a thin 0.5 points. For most workloads that number is within noise. On agentic coding the 2.4-point spread is not.

The Cost Layer

At $0.497 per successful LiveBench task, Terra costs $0.09 less than Sol’s $0.589 — a 16% saving that compounds over sustained agent runs. For a production agent executing thousands of tasks per day, Terra’s efficiency advantage at its specific strength category is the relevant number, not the headline composite.

Fable 5 Max Effort, for comparison, costs $1.573 per successful task. At its 46.9% agentic coding score — the weakest performance in the top tier — Fable 5 is more than three times the cost of Terra for agent workloads, with less than 70% of Terra’s category score.

The Selection Decision

LiveBench’s composite scores are weighted equally across all eight categories. Most agent deployments are not. Teams running production coding agents should evaluate models on the agentic coding dimension specifically — where Terra leads the current field — rather than headline composite scores that average in language, instruction-following, and mathematical reasoning.

The practical case for Sol over Terra exists: it is the stronger model for mixed workloads where coding agents run alongside structured document tasks, mathematical reasoning, and complex multi-modal instruction execution. If agentic coding is the primary workload, the current LiveBench data says Terra is the better buy.

That finding is not stable. OpenAI ships model updates across its Sol lineup without version bumps. The next update to either Terra or Sol could close or reverse the 2.4-point gap. Teams should re-run their own agentic benchmarks on each OpenAI update, rather than treating the LiveBench result as a permanent ranking.