GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Poolside's Laguna S 2.1: 118B MoE Hits 70.2% Terminal-Bench at $0.10 Per Million

Poolside released Laguna S 2.1 on July 21, an open-weight coding agent model that delivers competitive benchmark results at a price point well below the frontier tier.

Architecture and Benchmarks

Laguna S 2.1 uses a mixture-of-experts design: 118B total parameters, 8B active on any given forward pass. The architecture allows inference costs to stay close to a 7-10B dense model while the routing layer can draw on a larger effective parameter pool for harder tasks.

Key benchmark results:

BenchmarkLaguna S 2.1
Terminal-Bench 2.170.2%
DeepSWE40.4%

For context, the Terminal-Bench 2.1 leaderboard currently has GPT-5.5 at 84.7% (top cluster), Gemini 3.5 Flash at 76.2%, and Claude Sonnet 4.6 at roughly 74%. Laguna S 2.1 at 70.2% sits in the second tier — below the frontier cluster but well ahead of most open-weight models that can run affordably on standard hardware.

The DeepSWE score of 40.4% positions it similarly: below the frontier proprietary tier (Fable 5 leads at 80.3% on SWE-Bench Pro, GPT-5.5 in the low-to-mid 70s) but ahead of many models at its active-parameter count.

Pricing and Access

  • Input: $0.10/1M tokens
  • Output: $0.20/1M tokens
  • Context: 1M tokens
  • License: OpenMDW-1.1 (open-weight)

At $0.20/M output, Laguna S 2.1 is approximately 37x cheaper than Gemini 3.6 Flash ($7.50/M output), 125x cheaper than Claude Fable 5 ($25/M output), and about 3.5x cheaper than DeepSeek V4 Flash ($0.87/M after the permanent discount). For applications where 70% Terminal-Bench performance is sufficient, the cost argument is difficult to ignore.

The model is hosted via OpenRouter, where it currently has a single provider with no routing decisions. Poolside says it is intended for software engineering and agentic coding use cases.

What Laguna S 2.1 Is Not

The model is not positioned as a frontier replacement. Its 40.4% DeepSWE score is roughly half of what Fable 5 and GPT-5.5 achieve, and the Terminal-Bench number leaves a ~14-point gap to the leading cluster. For agentic workflows that require multi-step planning, context retention across long sessions, or high-stakes code modification, the frontier tier remains a different product.

The OpenMDW-1.1 license is not fully permissive — users are responsible for confirming appropriate use cases, and the model comes with an acceptable-use policy from Poolside.

Position in the Coding Agent Market

Laguna S 2.1 lands in a tier that is becoming increasingly competitive: open-weight models with MoE architectures in the 8-25B active parameter range. Kimi K2.7-Code, DeepSeek V4 Flash, and Ornith-1.0-397B all occupy similar territory. What distinguishes Laguna S 2.1 is the Terminal-Bench number relative to its cost: at $0.10/$0.20 per million, it undercuts nearly every comparable performer.

The release follows the trend of MoE models closing the gap to dense frontier models on coding-specific benchmarks while remaining economically viable for high-volume agent deployments.