Sakana AI Forms First Dedicated RSI Lab: Darwin Gödel Machine Doubled SWE Scores, SIFT Adds 11 Points for $25
Sakana AI, the Tokyo-based lab founded by ex-Google Brain researchers including David Ha, formally established a Recursive Self-Improvement (RSI) Lab today — the first dedicated research group at any independent AI organization explicitly tasked with building AI that redesigns its own development process.
The announcement lands in the same week Anthropic published data showing Claude now writes more than 80% of code merged into its own production systems. Sakana is making the same observation the explicit product strategy.
What RSI Means Here
Recursive self-improvement is AI improving its own capability, not just its outputs. The distinction matters: a model that writes better code is a tool; an agent that rewrites the scaffolding, prompts, and weights that determine how it writes code is something structurally different.
Sakana’s RSI Lab argues this capability is already partially real, pointing to a two-year publication record:
Darwin Gödel Machine — Agents that autonomously rewrite their own codebase. Sakana claims this doubled software-engineering benchmark performance without human-authored patches.
LLM-Squared (LLM²) — Developed with Oxford and Cambridge. An LLM inventing better ways to train LLMs through generational evolutionary loops. Produced DiscoPOP, a state-of-the-art preference optimization algorithm discovered and written entirely by an LLM.
ShinkaEvolve — Hyper-sample-efficient program evolution that builds novel loss functions for mixture-of-experts models.
ALE-Agent — Reinforcement learning agents that outperformed hundreds of human experts via self-learning.
The AI Scientist — End-to-end AI research automation, published in Nature.
SIFT: $25, 15 Hours, 11 Points
The most concrete near-term result comes from a companion paper published at ICLR 2026 (MIT and Sakana AI). SIFT (Self-Improvement via Fast Tree Search) identified the main bottleneck in recursive self-improvement: evaluation cost. Running a full benchmark to assess each candidate self-modification is expensive. SIFT replaces most evaluations with an LLM-as-a-judge, running the expensive benchmark only on the top-ranked patches.
Results on a 60-task SWE-bench Verified subset: starting from 51.7% with gpt-5-mini, SIFT reaches 61.7% in fewer than three self-improvement steps. Total cost: $25 in API spend, 15 CPU hours. That final score exceeds Kimi-K2-Thinking (the strongest open-source model at the time) and matches Claude Opus 4 on the same harness.
The judge model fidelity matters: gpt-5.2 as judge shows 34% Pearson correlation with actual benchmark performance; gpt-5-nano shows near-zero. Strong reasoning ability in the evaluator is the rate-limiting factor.
SIA: Both Levers Together
A related arxiv paper (SIA: Self-Improving AI with Harness and Weight Updates) pushes further by updating both the harness and model weights in a combined loop. Results across three domains:
- Legal (LawBench, 191-class Chinese charge classification): 13.5% → 70.1% accuracy. Harness-only peaked at 50%; weight updates via GRPO added another 20 points.
- GPU kernels (H100 Triton optimization): 91.9% runtime reduction from harness-only peak; 14x speedup over baseline.
- Biology (single-cell RNA denoising): 502% improvement over initial baseline; weight updates discovered a biological invariant (non-negative integer rounding) that scaffold iteration never generated across all iterations.
The key finding: the two levers operate in distinct change spaces. Harness updates change how the agent searches and acts; weight updates build domain intuition that prompts cannot encode. Neither saturates the other’s gains.
The Compute Angle
Sakana’s explicit positioning is against hyperscale. Their lab announcement argues RSI should be “a democratized public good” achievable on modest, sample-efficient compute — contrasting with the $50B+ compute clusters where frontier labs run training runs.
SIFT’s $25 result supports this framing for narrow domains. Whether it generalizes to arbitrary capability improvement across open-ended tasks is what the RSI Lab is now tasked to answer.
The lab is hiring frontier researchers in Tokyo, framing Japan’s compute constraints as a design forcing function rather than a disadvantage.