GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Gemini 3.7 Flash: DeepSWE Jumps 16 Points to 65.3%, Priced at Half of What 3.6 Flash Cost at Launch

Google released Gemini 3.7 Flash on August 13, three weeks after Gemini 3.6 Flash. This is not a new pretraining run. Google attributes the gains to algorithmic improvements to the core reasoning foundation — the same architecture, re-tuned.

The gains are substantial enough to matter anyway.

Benchmark Numbers

Benchmark3.6 Flash3.7 FlashDelta
FrontierCode 1.1 Main34.4%43.6%+9.2 pts
DeepSWE v1.149.0%65.3%+16.3 pts
WebDev Arena Elo15381588+50 pts
GDP.pdf22.0%34.0%+12 pts
AutomationBench17.0%30.4%+13.4 pts

DeepSWE is the headline number. A 16-point jump in software engineering resolution is significant — it’s the difference between a model that handles routine tickets and one that works on real codebases. Per Google’s internal benchmarks, 3.7 Flash now sits ahead of both Claude Sonnet 5 and GPT-5.6 Terra on DeepSWE.

The AutomationBench gain (+13.4 points) is the less-reported signal worth watching. AutomationBench tests real-world business workflow completion — the agentic layer that most enterprise customers care about most. That score nearly doubling suggests the reasoning improvements translate to multi-step, real-world task chains, not just code puzzles.

Pricing

Launch pricing is $0.75 per million input tokens and $3.75 per million output tokens. That holds through the end of the year. Both 3.6 Flash and 3.7 Flash now sit at the same price point — meaning 3.6 Flash got a retroactive cut, and 3.7 Flash comes in at half what 3.6 Flash originally cost at launch.

At $0.75/M input, this is the lowest entry point for a Google Flash tier model since the series began.

Availability

3.7 Flash is available now through the Gemini API, AI Studio, and Antigravity. Google is also rolling the model into Gemini Spark — the reasoning mode available to AI Pro and Ultra subscribers across 160+ countries.

What This Is and Isn’t

This is an algorithmic improvement release, not a frontier push. Claude Opus 5 at 97% SWE-Bench Verified and GPT-5.6 Sol at 96.2% sit well above 3.7 Flash’s territory. What Google is doing with Flash is different: shipping a capable-enough workhorse at a price point that makes agentic workflows economically viable at scale.

For teams running agent pipelines at volume, $0.75/M input changes the math. A 65.3% DeepSWE score at that price is a serious production option. Gemini 3.5 Pro — Google’s still-absent larger flagship — remains the unshipped promise. Flash keeps shipping in its place.