GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Gemini 3.5 Flash Leads Terminal-Bench 2.1 at 76.2% — 4x Faster Than Rivals at Under Half the Cost

Google shipped Gemini 3.5 Flash today at I/O 2026, opening a new model family that the company says is designed from the ground up for agentic and coding workloads. The headline numbers: 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA, and 83.6% on MCP Atlas. All three put 3.5 Flash ahead of Gemini 3.1 Pro — a flagship-tier model — at Flash-tier speed and price.

Google claims the model runs four times faster than other frontier models on output tokens per second. At $0.75/M input and $4.50/M output (standard context, Google Cloud Vertex AI), it undercuts comparable frontier models by more than half on cost. VentureBeat cites Google’s internal estimate that enterprise customers running high-volume agentic workloads could save over $1 billion per year by switching from current frontier options.

Benchmark Breakdown

BenchmarkGemini 3.5 FlashPrior best (3.1 Pro)
Terminal-Bench 2.176.2%~80.2% (Gemini 3.1 Pro, TongAgents)
GDPval-AA Elo1656Comparison pending
MCP Atlas83.6%—
CharXiv Reasoning84.2%—

One clarification on Terminal-Bench: the 3.5 Flash score is a direct agent submission, while the Terminal-Bench 2.0 leaderboard shows vix+Claude Opus 4.7 at 90.2% and GPT-5.5 systems above 80%. The 3.5 Flash score is Google’s unscaffolded submission and lands the model in a competitive position among Flash-class models.

Antigravity Integration

The model ships with tight integration into Google Antigravity, Google’s agent-first development platform. Under Antigravity, 3.5 Flash can spawn and coordinate collaborative subagents for multi-step workflows — file categorization, codebase maintenance, financial document preparation. Google demonstrated a multi-agent pipeline where 3.5 Flash autonomously renamed and categorized large sets of unstructured assets using dynamic criteria, running multiple subagents in parallel.

The speed profile is key here. Flash-class models occupy the slot where agent loops need to run at scale without the per-token cost of a flagship model eating into economics. 3.5 Flash makes a stronger case for that role than prior Flash generations.

Availability

Gemini 3.5 Flash is live today across:

  • Gemini app (billions of users)
  • AI Mode in Google Search
  • Google Antigravity agent platform
  • Gemini API via Google AI Studio and Android Studio
  • Gemini Enterprise Agent Platform and Gemini Enterprise

Pricing

Context tierInputOutput
Standard (<200K tokens)$0.75/M$4.50/M
Long context (>200K tokens)$1.50/M$9.00/M
Extended long context$2.70/M$16.20/M

What’s Next

Gemini 3.5 Pro is already in internal use at Google and is scheduled for external rollout next month. The Pro variant is expected to compete directly at the top of the intelligence index alongside GPT-5.5 and Claude Opus 4.7. Based on the Flash-to-Pro gap in prior Gemini generations, Pro should post materially higher scores on reasoning and math benchmarks while carrying a higher price tag.

Google is entering I/O with a model that competes on benchmark performance, undercuts on cost, and leads on raw throughput. The competitive argument for 3.5 Flash is strongest in high-volume agentic applications where running 10,000 turns per day at GPT-5.5 pricing is not viable.