GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Cognition's SWE-2 Claims Frontier Parity at 64% Lower Cost - Then Posts 27% on Terminal-Bench 4

Cognition launched SWE-2 on September 10, calling it the most advanced coding model yet from the team that built Devin. The headline number is 50.0% on FrontierCode 1.1 Main, landing within one point of Claude Fable 5.1 (50.9%) while priced 64% cheaper. On DeepSWE 1.1, SWE-2 (73.0%) outperforms Fable 5.1 (67.4%) and every other model in the comparison except GPT-6 Astra (74.1%).

BenchmarkSWE-2Fable 5.1GPT-6 AstraGrok 4.6Kimi K3
FrontierCode 1.1 Main50.0%50.9%53.3%48.0%44.2%
DeepSWE 1.173.0%67.4%74.1%67.5%68.5%
Terminal-Bench 2.192.8%————

Against SWE-1.7 (42.0% on FrontierCode 1.1 Main), the 8-point improvement is real and significant. The cost argument is also genuine. If those two benchmark columns held up on their own, this would be a clean Pareto story.

The Number Cognition Did Not Headline

Terminal-Bench 4: 27.3%.

Terminal-Bench 2.1 and Terminal-Bench 4 are successive versions of the same agentic software engineering benchmark. SWE-2 scored 92.8% on the older version — the highest any model has posted on Terminal-Bench 2.1. On the current version, it scores 27.3%. That is a 65-point collapse between two releases of the same benchmark measuring the same underlying capability class.

For comparison: models that genuinely generalise tend to move within a narrower band across benchmark versions. A 65-point gap is not measurement noise.

The most straightforward explanation is test-set contamination. Terminal-Bench 2.1 has been public long enough for its task distributions to enter training pipelines, deliberately or through web crawl. A model that has absorbed those distributions can score extremely high without having learned the underlying skill. Terminal-Bench 4, with newer and presumably cleaner tasks, strips that advantage away.

HN reviewers flagged the delta immediately after launch. Cognition has not publicly addressed it.

Cognition’s Structural Position

Cognition now owns Windsurf, the coding IDE, after the OpenAI acquisition collapsed earlier this year. That gives Cognition access to real software engineering sessions at scale, which should be a meaningful training data asset. The irony is that a company with that much real-world coding data is the one posting a benchmark result that looks like it may have been over-indexed on a synthetic test set.

The FrontierCode and DeepSWE results suggest SWE-2 has real capability. The benchmark selection in the launch post, which leads with Terminal-Bench 2.1 and omits Terminal-Bench 4, suggests Cognition knew the newer version was a problem.

The Practical Read

SWE-2 is likely a capable model at a competitive price. The FrontierCode and DeepSWE numbers are from benchmarks that are harder to game at scale and those scores look defensible. For teams evaluating coding models on cost-performance grounds, the 64% price advantage over Fable 5.1 with near-parity on FrontierCode is worth testing against real workloads.

The Terminal-Bench 2.1 result should be disregarded until Cognition either explains the Terminal-Bench 4 score or posts an independent replication on the newer version. A 65-point gap between benchmark versions is a disclosure problem, not just a performance problem.