Sakana's Fugu Ultra Posts 73.7% SWE-Bench Pro — Collective AI Beats Frontier Monoliths, 10 Days After Fable Ban
Japanese AI startup Sakana AI launched Fugu and Fugu Ultra on June 22, a multi-model orchestration system that posts frontier-level benchmark numbers without owning a frontier model. Fugu Ultra scores 73.7% on SWE-Bench Pro, 82.1% on TerminalBench 2.1, and 93.2% on LiveCodeBench — all ahead of Anthropic’s Claude Opus 4.8 and OpenAI’s GPT-5.5. The launch came exactly ten days after the US government ordered Anthropic to pull Fable 5 and Mythos offline globally, and Sakana did not waste the timing.
What It Actually Is
Fugu is not a model. It is a routing system that coordinates a swappable pool of publicly available language models through a single OpenAI-compatible API. Users submit tasks; Fugu decides internally which models to invoke, in what order, and how to synthesise the results. The underlying pool excludes Fable 5 and Mythos, since neither is publicly accessible. If those models return to availability, Sakana says scores would be higher still.
Two tiers: Fugu for coding, chat, and everyday tasks. Fugu Ultra for complex work including AI research, cybersecurity analysis, and multi-step patent investigation. Both ship now, with subscription plans and usage-based billing.
Benchmark Numbers
Sakana published a full comparison against Opus 4.8, Gemini 3.1 Pro, and GPT-5.5. The table is the whole story:
| Benchmark | Fugu Ultra | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
|---|---|---|---|---|
| SWE-Bench Pro | 73.7% | 69.2% | 58.6% | 54.2% |
| TerminalBench 2.1 | 82.1% | 74.6% | 78.2% | 70.3% |
| LiveCodeBench | 93.2% | 87.8% | 85.3% | 88.5% |
| LiveCodeBench Pro | 90.8% | 84.8% | 88.4% | 82.9% |
| GPQA-D | 95.5% | 92.0% | 93.6% | 94.3% |
| Humanity’s Last Exam | 50.0% | 49.8% | 41.4% | 44.4% |
| CharXiv Reasoning | 86.6% | 84.2% | 84.1% | 83.3% |
Fugu Ultra beats every accessible frontier model on six of seven benchmarks. SWE-Bench Pro is the clearest signal: 73.7% against Opus 4.8 at 69.2% and GPT-5.5 at 58.6% is not marginal.
Important caveat: baseline scores for non-Fugu models come from their respective providers. Fugu’s scores come from Sakana’s own evaluation using mini-SWE-agent scaffolding. Independent replication hasn’t landed yet.
There are also benchmarks where Fugu doesn’t lead. Fable 5, while restricted, published 80.0% on SWE-Bench Pro and 53.3% on Humanity’s Last Exam — both ahead of Fugu Ultra. The orchestrator closes the gap, it doesn’t eliminate it.
The Architecture Debate
The launch immediately split the developer community on Reddit and elsewhere. The core question: can a system that routes across other models claim to be a frontier model itself? Critics note that Fugu is buying capability, not building it, and that benchmark comparisons are meaningless if the competing models are also available in Fugu’s pool.
Defenders point out that frontier labs have blurred this line themselves: GPT-5.5 internally routes across multiple specialised heads, and Mixture-of-Experts architectures are effectively routing systems anyway. If the output is indistinguishable, the internal architecture may not matter to enterprise buyers.
For Sakana, the framing is simpler: reduced vendor risk. Every closed-model dependency is a single point of failure after June 12.
Strategic Timing
Sakana was founded in 2024 by two former Google DeepMind researchers and was valued at $2.6 billion in a Series B closed in late 2025. The June 22 launch followed a June 15 release of Marlin, an eight-hour autonomous research agent targeting B2B workflows. Both products share the same underlying routing infrastructure.
The Fable 5 and Mythos export ban created a demand signal that Sakana had been building toward for months. Enterprises that spent a year integrating Anthropic’s models suddenly needed a hedged alternative. Fugu’s pitch is exactly that: frontier-level output with no single-vendor dependency and a pool that swaps automatically as models are added or restricted.
SpaceX has signed Reflection AI to a $6.3 billion compute deal to build open-weight frontier alternatives. Cohere’s Command A+ shipped open-weights under Apache 2.0. The Fable 5 shutdown appears to be restructuring how enterprise AI thinks about model concentration risk — and Sakana is betting that restructuring is permanent.