Gemini 3.5 Flash Leads Terminal-Bench 2.1 at 76.2% — 4x Faster Than Rivals at Under Half the Cost
Google shipped Gemini 3.5 Flash today at I/O 2026, opening a new model family that the company says is designed from the ground up for agentic and coding workloads. The headline numbers: 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA, and 83.6% on MCP Atlas. All three put 3.5 Flash ahead of Gemini 3.1 Pro — a flagship-tier model — at Flash-tier speed and price.
Google claims the model runs four times faster than other frontier models on output tokens per second. At $0.75/M input and $4.50/M output (standard context, Google Cloud Vertex AI), it undercuts comparable frontier models by more than half on cost. VentureBeat cites Google’s internal estimate that enterprise customers running high-volume agentic workloads could save over $1 billion per year by switching from current frontier options.
Benchmark Breakdown
| Benchmark | Gemini 3.5 Flash | Prior best (3.1 Pro) |
|---|---|---|
| Terminal-Bench 2.1 | 76.2% | ~80.2% (Gemini 3.1 Pro, TongAgents) |
| GDPval-AA Elo | 1656 | Comparison pending |
| MCP Atlas | 83.6% | — |
| CharXiv Reasoning | 84.2% | — |
One clarification on Terminal-Bench: the 3.5 Flash score is a direct agent submission, while the Terminal-Bench 2.0 leaderboard shows vix+Claude Opus 4.7 at 90.2% and GPT-5.5 systems above 80%. The 3.5 Flash score is Google’s unscaffolded submission and lands the model in a competitive position among Flash-class models.
Antigravity Integration
The model ships with tight integration into Google Antigravity, Google’s agent-first development platform. Under Antigravity, 3.5 Flash can spawn and coordinate collaborative subagents for multi-step workflows — file categorization, codebase maintenance, financial document preparation. Google demonstrated a multi-agent pipeline where 3.5 Flash autonomously renamed and categorized large sets of unstructured assets using dynamic criteria, running multiple subagents in parallel.
The speed profile is key here. Flash-class models occupy the slot where agent loops need to run at scale without the per-token cost of a flagship model eating into economics. 3.5 Flash makes a stronger case for that role than prior Flash generations.
Availability
Gemini 3.5 Flash is live today across:
- Gemini app (billions of users)
- AI Mode in Google Search
- Google Antigravity agent platform
- Gemini API via Google AI Studio and Android Studio
- Gemini Enterprise Agent Platform and Gemini Enterprise
Pricing
| Context tier | Input | Output |
|---|---|---|
| Standard (<200K tokens) | $0.75/M | $4.50/M |
| Long context (>200K tokens) | $1.50/M | $9.00/M |
| Extended long context | $2.70/M | $16.20/M |
What’s Next
Gemini 3.5 Pro is already in internal use at Google and is scheduled for external rollout next month. The Pro variant is expected to compete directly at the top of the intelligence index alongside GPT-5.5 and Claude Opus 4.7. Based on the Flash-to-Pro gap in prior Gemini generations, Pro should post materially higher scores on reasoning and math benchmarks while carrying a higher price tag.
Google is entering I/O with a model that competes on benchmark performance, undercuts on cost, and leads on raw throughput. The competitive argument for 3.5 Flash is strongest in high-volume agentic applications where running 10,000 turns per day at GPT-5.5 pricing is not viable.