Poolside's Laguna S 2.1: 118B MoE Hits 70.2% Terminal-Bench at $0.10 Per Million
Poolside released Laguna S 2.1 on July 21, an open-weight coding agent model that delivers competitive benchmark results at a price point well below the frontier tier.
Architecture and Benchmarks
Laguna S 2.1 uses a mixture-of-experts design: 118B total parameters, 8B active on any given forward pass. The architecture allows inference costs to stay close to a 7-10B dense model while the routing layer can draw on a larger effective parameter pool for harder tasks.
Key benchmark results:
| Benchmark | Laguna S 2.1 |
|---|---|
| Terminal-Bench 2.1 | 70.2% |
| DeepSWE | 40.4% |
For context, the Terminal-Bench 2.1 leaderboard currently has GPT-5.5 at 84.7% (top cluster), Gemini 3.5 Flash at 76.2%, and Claude Sonnet 4.6 at roughly 74%. Laguna S 2.1 at 70.2% sits in the second tier — below the frontier cluster but well ahead of most open-weight models that can run affordably on standard hardware.
The DeepSWE score of 40.4% positions it similarly: below the frontier proprietary tier (Fable 5 leads at 80.3% on SWE-Bench Pro, GPT-5.5 in the low-to-mid 70s) but ahead of many models at its active-parameter count.
Pricing and Access
- Input: $0.10/1M tokens
- Output: $0.20/1M tokens
- Context: 1M tokens
- License: OpenMDW-1.1 (open-weight)
At $0.20/M output, Laguna S 2.1 is approximately 37x cheaper than Gemini 3.6 Flash ($7.50/M output), 125x cheaper than Claude Fable 5 ($25/M output), and about 3.5x cheaper than DeepSeek V4 Flash ($0.87/M after the permanent discount). For applications where 70% Terminal-Bench performance is sufficient, the cost argument is difficult to ignore.
The model is hosted via OpenRouter, where it currently has a single provider with no routing decisions. Poolside says it is intended for software engineering and agentic coding use cases.
What Laguna S 2.1 Is Not
The model is not positioned as a frontier replacement. Its 40.4% DeepSWE score is roughly half of what Fable 5 and GPT-5.5 achieve, and the Terminal-Bench number leaves a ~14-point gap to the leading cluster. For agentic workflows that require multi-step planning, context retention across long sessions, or high-stakes code modification, the frontier tier remains a different product.
The OpenMDW-1.1 license is not fully permissive — users are responsible for confirming appropriate use cases, and the model comes with an acceptable-use policy from Poolside.
Position in the Coding Agent Market
Laguna S 2.1 lands in a tier that is becoming increasingly competitive: open-weight models with MoE architectures in the 8-25B active parameter range. Kimi K2.7-Code, DeepSeek V4 Flash, and Ornith-1.0-397B all occupy similar territory. What distinguishes Laguna S 2.1 is the Terminal-Bench number relative to its cost: at $0.10/$0.20 per million, it undercuts nearly every comparable performer.
The release follows the trend of MoE models closing the gap to dense frontier models on coding-specific benchmarks while remaining economically viable for high-volume agent deployments.