GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Gemini 3.5 Flash Lite Targets the Sub-Agent Layer at $0.30/$2.50 Per Million

Google released Gemini 3.5 Flash Lite on July 21, positioning it explicitly as infrastructure for the sub-agent layer rather than a general-purpose capable model. The pitch: you do not need a frontier model to execute a focused task inside a larger pipeline — and you definitely do not want to pay frontier prices for it.

Specs

  • Input: $0.30 per million tokens
  • Output: $2.50 per million tokens
  • Speed: 174 tokens per second on Google AI Studio, 56 tps on Vertex
  • Context: 1M tokens
  • Cache write: $1.25/M; cache hit: $0.10/M (90% discount)
  • Modalities: text and image input, text output
  • Released: July 21, 2026

The cache discount is the number that matters for agent deployments. Sub-agents in a multi-agent system typically receive the same system prompt and substantial shared context on every call. At $0.10/M on cache hits, the effective per-token cost for repeated context drops by 90%. OpenRouter’s 30-day weighted average for actual customer pricing already reflects this: $0.128/M input, $2.34/M output after caching.

The Sub-Agent Positioning

Google is explicit about the use case: “suited for subagents that execute focused tasks within complex, multi-agent workflows.” Flash Lite is not meant to orchestrate or reason over long chains — it is meant to receive a scoped instruction and complete it, fast, cheaply, and reliably.

This is the segment that Gemini 3.5 Flash Lite is entering. The roster of models competing here has grown significantly: Gemini 3.1 Flash Lite at $0.25/M input (covered earlier as Google’s half-price Flash with thinking levels), DeepSeek V4 Flash, and the lower tiers of the GPT-5.6 family. Flash Lite’s 174 tps puts it competitive on throughput.

Why This Matters

Multi-agent systems have a cost structure problem. The orchestrator model drives quality, but the sub-agents drive volume — and if sub-agents run on frontier pricing, even moderate task complexity makes costs untenable. The June Goldman analysis projected token use multiplying 24x by 2030 as agentic workflows scale. The economic pressure to route sub-agent calls to the cheapest acceptable model is already reshaping how teams architect pipelines.

Flash Lite at $2.50/M output is cheaper than Gemini 3.5 Flash proper, cheaper than any Anthropic model above Haiku, and significantly cheaper than GPT-5.5 or GPT-5.6. Whether it holds quality for sub-agent tasks depends on the domain — but the pricing gives it a clear position at the bottom of the capable tier.

Google’s timing is deliberate. Cloudflare shipped five agent products in three days last month. AWS AgentCore added persistent filesystem and shell execution. The infrastructure layer is hardening around multi-agent execution, and every infrastructure play needs a cheap, fast, reliable model at the leaf node. Flash Lite is Google’s bid for that position.