Gemini 3.7 Flash: DeepSWE Jumps 16 Points to 65.3%, Priced at Half of What 3.6 Flash Cost at Launch
Google released Gemini 3.7 Flash on August 13, three weeks after Gemini 3.6 Flash. This is not a new pretraining run. Google attributes the gains to algorithmic improvements to the core reasoning foundation — the same architecture, re-tuned.
The gains are substantial enough to matter anyway.
Benchmark Numbers
| Benchmark | 3.6 Flash | 3.7 Flash | Delta |
|---|---|---|---|
| FrontierCode 1.1 Main | 34.4% | 43.6% | +9.2 pts |
| DeepSWE v1.1 | 49.0% | 65.3% | +16.3 pts |
| WebDev Arena Elo | 1538 | 1588 | +50 pts |
| GDP.pdf | 22.0% | 34.0% | +12 pts |
| AutomationBench | 17.0% | 30.4% | +13.4 pts |
DeepSWE is the headline number. A 16-point jump in software engineering resolution is significant — it’s the difference between a model that handles routine tickets and one that works on real codebases. Per Google’s internal benchmarks, 3.7 Flash now sits ahead of both Claude Sonnet 5 and GPT-5.6 Terra on DeepSWE.
The AutomationBench gain (+13.4 points) is the less-reported signal worth watching. AutomationBench tests real-world business workflow completion — the agentic layer that most enterprise customers care about most. That score nearly doubling suggests the reasoning improvements translate to multi-step, real-world task chains, not just code puzzles.
Pricing
Launch pricing is $0.75 per million input tokens and $3.75 per million output tokens. That holds through the end of the year. Both 3.6 Flash and 3.7 Flash now sit at the same price point — meaning 3.6 Flash got a retroactive cut, and 3.7 Flash comes in at half what 3.6 Flash originally cost at launch.
At $0.75/M input, this is the lowest entry point for a Google Flash tier model since the series began.
Availability
3.7 Flash is available now through the Gemini API, AI Studio, and Antigravity. Google is also rolling the model into Gemini Spark — the reasoning mode available to AI Pro and Ultra subscribers across 160+ countries.
What This Is and Isn’t
This is an algorithmic improvement release, not a frontier push. Claude Opus 5 at 97% SWE-Bench Verified and GPT-5.6 Sol at 96.2% sit well above 3.7 Flash’s territory. What Google is doing with Flash is different: shipping a capable-enough workhorse at a price point that makes agentic workflows economically viable at scale.
For teams running agent pipelines at volume, $0.75/M input changes the math. A 65.3% DeepSWE score at that price is a serious production option. Gemini 3.5 Pro — Google’s still-absent larger flagship — remains the unshipped promise. Flash keeps shipping in its place.