Gemini 3.6 Flash is Google’s July 21, 2026 workhorse model — the successor to Gemini 3.5 Flash and a deliberate move to make agentic coding cheaper without sacrificing quality. The pitch is efficiency: 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, takes fewer reasoning steps and tool calls, and costs less per output token ($7.50 vs $9.00).
It launched alongside two siblings: Gemini 3.5 Flash-Lite ($0.30/$2.50, the budget tier) and Gemini 3.5 Flash Cyber (a security-focused variant, limited to governments and trusted partners). Notably absent from the launch: Gemini 3.5 Pro, which Google has been teasing since May but has reportedly struggled to meet internal performance goals.
Specs at a glance
| Spec | Value |
|---|---|
| Model ID | gemini-3.6-flash |
| Context | 1M tokens (1,048,576 input / 65,536 output) |
| Modalities | Text, image, video, audio, PDF in → text out |
| Knowledge cutoff | March 2026 (some domains may lag to Jan 2025) |
| Thinking | Supported |
| Computer use | Preview |
| Tools | Function calling, code execution, search grounding, file search, structured outputs |
Pricing
| 3.6 Flash | 3.5 Flash | Change | |
|---|---|---|---|
| Input / 1M | $1.50 | $1.50 | — |
| Output / 1M | $7.50 | $9.00 | −17% |
| Cached input / 1M | $0.375 | $0.375 | — |
The input price is unchanged; the output price drops by $1.50 per million tokens. For agentic workloads where output dominates (multi-step reasoning, code generation), the effective cost reduction is larger than 17% because the model also produces fewer tokens per task.
Flex inference and Priority inference are both supported. Batch API is available.
Benchmarks
Google frames 3.6 Flash’s gains around efficiency rather than peak score — the model does better work with fewer tokens. Key vendor-reported results:
- DeepSWE by Datacurve: up to 65% fewer output tokens than 3.5 Flash on the same tasks
- Artificial Analysis Index: 17% fewer output tokens overall, with a 12-point gain on coding tasks
- Computer use: double-digit improvement over 3.5 Flash (Preview)
- Knowledge work: measurable gains on coding, knowledge and multimodal benchmarks
Google has not published a full side-by-side benchmark table at launch. Treat efficiency claims as vendor-reported; prefer independent AA/community replications.
Also shipped: 3.5 Flash-Lite and Flash Cyber
Gemini 3.5 Flash-Lite ($0.30 input / $2.50 output per 1M) is Google’s most cost-effective model — slightly pricier than the previous Flash-Lite ($0.25/$1.50) but with benchmark scores that in several cases beat the full-size 3.5 Flash. It is available to everyone and rolling out in Google Search.
Gemini 3.5 Flash Cyber is fine-tuned for finding and fixing cybersecurity vulnerabilities. It is exclusively available to governments and trusted partners as part of a limited-access pilot. No public pricing has been disclosed.
Availability
| Surface | Status |
|---|---|
| Gemini API (Google AI Studio) | Available (model gemini-3.6-flash) |
| Android Studio | Available |
| Google Antigravity | Available |
| Gemini Enterprise Agent Platform | Available |
| Gemini app (consumer) | Available |
How it compares
- vs 3.5 Flash: cheaper output ($7.50 vs $9.00), more efficient (17% fewer tokens), better benchmarks across coding/knowledge/computer-use. Direct upgrade.
- vs Gemini 3.1 Pro: 3.6 Flash is positioned as the agentic workhorse; 3.1 Pro remains the intelligence flagship at ~$2/$12. For most agent loops, 3.6 Flash is the better cost pick.
- vs Claude Sonnet 5: Sonnet 5 at $2/$10 intro is close on input pricing; 3.6 Flash wins on output ($7.50 vs $10) and native multimodal input (video, audio, PDF). Sonnet 5 leads on agentic coding benchmarks.
- vs GPT-5.6 Luna: Luna ($1/$6) is cheaper and lighter; 3.6 Flash ($1.50/$7.50) is more capable and natively multimodal.
The missing flagship
The launch is notable for what Google didn’t ship: Gemini 3.5 Pro. Google teased it when 3.5 Flash launched in May (“already being used internally, and we look forward to rolling it out next month”), but Bloomberg reported internal delays in July. Product lead Logan Kilpatrick confirmed on launch day that 3.5 Pro is “currently testing with partners” and that the team has started pre-training for Gemini 4. The Flash tier keeps getting better; the Pro tier remains 3.1.
Best for
- High-volume agentic coding loops where token cost dominates
- Multi-step tool-calling workflows and agent chains
- Multimodal document/video/audio processing at scale
- Cost-conscious alternative to frontier models for knowledge work
Pros & cons
Strengths
- 17% fewer output tokens than 3.5 Flash on the AA Index — fewer reasoning steps, lower per-task cost
- Output price drops to $7.50 / 1M (from $9.00) — cheapest capable frontier Flash model
- Double-digit gains over 3.5 Flash on DeepSWE, computer-use and knowledge benchmarks
- Natively multimodal — text, image, video, audio and PDF input
Limitations
- No image generation or Live API support
- Computer use still in Preview
- Knowledge cutoff March 2026 (some domains lag to Jan 2025, per model card)
- Still no Gemini 3.5 Pro — 3.1 Pro remains the flagship
Frequently asked questions
When was Gemini 3.6 Flash released?
Google DeepMind released Gemini 3.6 Flash on July 21, 2026, alongside Gemini 3.5 Flash-Lite (budget tier) and Gemini 3.5 Flash Cyber (security-focused, limited access). 3.6 Flash replaces 3.5 Flash as Google's mid-tier workhorse.
How much does Gemini 3.6 Flash cost?
$1.50 per 1M input tokens and $7.50 per 1M output tokens. Input price matches 3.5 Flash; output drops from $9.00 (~17% cut). Flex inference is available at a further discount. Cached input starts at $0.375 / 1M.
What is the API model ID for Gemini 3.6 Flash?
Use `gemini-3.6-flash` in the Gemini API. It's available in Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise app and Agent Platform.
Is Gemini 3.6 Flash better than 3.5 Flash?
Yes, on published benchmarks. Google reports double-digit gains on DeepSWE, computer use and knowledge work, while using 17% fewer output tokens on the AA Index. It's also cheaper per output token ($7.50 vs $9.00).
What happened to Gemini 3.5 Pro?
As of the 3.6 Flash launch, Google has not released 3.5 Pro. TechCrunch reported Google was facing internal delays. Product lead Logan Kilpatrick said the team is testing it with partners and hopes to "land soon," and noted pre-training for Gemini 4 has begun.
Sources
- Google Blog — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- Google DeepMind — Gemini 3.6 Flash Model Card
- Google AI for Developers — Gemini 3.6 Flash
- TechCrunch — Google releases three new Gemini models — but no 3.5 Pro
- Tech Insider — Gemini 3.6 Flash Debuts: 17% Cheaper, 12-Point Gain