Large Language Models

Gemini 3.6 Flash

Google's July 2026 workhorse model — 17% fewer output tokens than 3.5 Flash, cheaper at $1.50/$7.50, optimized for agentic coding and knowledge work.

Gemini 3.6 Flash is Google’s July 21, 2026 workhorse model — the successor to Gemini 3.5 Flash and a deliberate move to make agentic coding cheaper without sacrificing quality. The pitch is efficiency: 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, takes fewer reasoning steps and tool calls, and costs less per output token ($7.50 vs $9.00).

It launched alongside two siblings: Gemini 3.5 Flash-Lite ($0.30/$2.50, the budget tier) and Gemini 3.5 Flash Cyber (a security-focused variant, limited to governments and trusted partners). Notably absent from the launch: Gemini 3.5 Pro, which Google has been teasing since May but has reportedly struggled to meet internal performance goals.

Specs at a glance

SpecValue
Model IDgemini-3.6-flash
Context1M tokens (1,048,576 input / 65,536 output)
ModalitiesText, image, video, audio, PDF in → text out
Knowledge cutoffMarch 2026 (some domains may lag to Jan 2025)
ThinkingSupported
Computer usePreview
ToolsFunction calling, code execution, search grounding, file search, structured outputs

Pricing

3.6 Flash3.5 FlashChange
Input / 1M$1.50$1.50
Output / 1M$7.50$9.00−17%
Cached input / 1M$0.375$0.375

The input price is unchanged; the output price drops by $1.50 per million tokens. For agentic workloads where output dominates (multi-step reasoning, code generation), the effective cost reduction is larger than 17% because the model also produces fewer tokens per task.

Flex inference and Priority inference are both supported. Batch API is available.

Benchmarks

Google frames 3.6 Flash’s gains around efficiency rather than peak score — the model does better work with fewer tokens. Key vendor-reported results:

  • DeepSWE by Datacurve: up to 65% fewer output tokens than 3.5 Flash on the same tasks
  • Artificial Analysis Index: 17% fewer output tokens overall, with a 12-point gain on coding tasks
  • Computer use: double-digit improvement over 3.5 Flash (Preview)
  • Knowledge work: measurable gains on coding, knowledge and multimodal benchmarks

Google has not published a full side-by-side benchmark table at launch. Treat efficiency claims as vendor-reported; prefer independent AA/community replications.

Also shipped: 3.5 Flash-Lite and Flash Cyber

Gemini 3.5 Flash-Lite ($0.30 input / $2.50 output per 1M) is Google’s most cost-effective model — slightly pricier than the previous Flash-Lite ($0.25/$1.50) but with benchmark scores that in several cases beat the full-size 3.5 Flash. It is available to everyone and rolling out in Google Search.

Gemini 3.5 Flash Cyber is fine-tuned for finding and fixing cybersecurity vulnerabilities. It is exclusively available to governments and trusted partners as part of a limited-access pilot. No public pricing has been disclosed.

Availability

SurfaceStatus
Gemini API (Google AI Studio)Available (model gemini-3.6-flash)
Android StudioAvailable
Google AntigravityAvailable
Gemini Enterprise Agent PlatformAvailable
Gemini app (consumer)Available

How it compares

  • vs 3.5 Flash: cheaper output ($7.50 vs $9.00), more efficient (17% fewer tokens), better benchmarks across coding/knowledge/computer-use. Direct upgrade.
  • vs Gemini 3.1 Pro: 3.6 Flash is positioned as the agentic workhorse; 3.1 Pro remains the intelligence flagship at ~$2/$12. For most agent loops, 3.6 Flash is the better cost pick.
  • vs Claude Sonnet 5: Sonnet 5 at $2/$10 intro is close on input pricing; 3.6 Flash wins on output ($7.50 vs $10) and native multimodal input (video, audio, PDF). Sonnet 5 leads on agentic coding benchmarks.
  • vs GPT-5.6 Luna: Luna ($1/$6) is cheaper and lighter; 3.6 Flash ($1.50/$7.50) is more capable and natively multimodal.

The missing flagship

The launch is notable for what Google didn’t ship: Gemini 3.5 Pro. Google teased it when 3.5 Flash launched in May (“already being used internally, and we look forward to rolling it out next month”), but Bloomberg reported internal delays in July. Product lead Logan Kilpatrick confirmed on launch day that 3.5 Pro is “currently testing with partners” and that the team has started pre-training for Gemini 4. The Flash tier keeps getting better; the Pro tier remains 3.1.

Best for

  • High-volume agentic coding loops where token cost dominates
  • Multi-step tool-calling workflows and agent chains
  • Multimodal document/video/audio processing at scale
  • Cost-conscious alternative to frontier models for knowledge work

Pros & cons

Strengths

  • 17% fewer output tokens than 3.5 Flash on the AA Index — fewer reasoning steps, lower per-task cost
  • Output price drops to $7.50 / 1M (from $9.00) — cheapest capable frontier Flash model
  • Double-digit gains over 3.5 Flash on DeepSWE, computer-use and knowledge benchmarks
  • Natively multimodal — text, image, video, audio and PDF input

Limitations

  • No image generation or Live API support
  • Computer use still in Preview
  • Knowledge cutoff March 2026 (some domains lag to Jan 2025, per model card)
  • Still no Gemini 3.5 Pro — 3.1 Pro remains the flagship

Frequently asked questions

When was Gemini 3.6 Flash released?

Google DeepMind released Gemini 3.6 Flash on July 21, 2026, alongside Gemini 3.5 Flash-Lite (budget tier) and Gemini 3.5 Flash Cyber (security-focused, limited access). 3.6 Flash replaces 3.5 Flash as Google's mid-tier workhorse.

How much does Gemini 3.6 Flash cost?

$1.50 per 1M input tokens and $7.50 per 1M output tokens. Input price matches 3.5 Flash; output drops from $9.00 (~17% cut). Flex inference is available at a further discount. Cached input starts at $0.375 / 1M.

What is the API model ID for Gemini 3.6 Flash?

Use `gemini-3.6-flash` in the Gemini API. It's available in Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise app and Agent Platform.

Is Gemini 3.6 Flash better than 3.5 Flash?

Yes, on published benchmarks. Google reports double-digit gains on DeepSWE, computer use and knowledge work, while using 17% fewer output tokens on the AA Index. It's also cheaper per output token ($7.50 vs $9.00).

What happened to Gemini 3.5 Pro?

As of the 3.6 Flash launch, Google has not released 3.5 Pro. TechCrunch reported Google was facing internal delays. Product lead Logan Kilpatrick said the team is testing it with partners and hopes to "land soon," and noted pre-training for Gemini 4 has begun.

Sources

Last updated: 2026-07-29 · Specs and pricing change fast — verify on the vendor's site before relying on them.