Gemini 3.6 Flash thumbnail with the headline “GEMINI 3.6 FLASH — IS IT WORTH IT?” beside a glowing blue and purple Gemini logo on a dark cinematic background.

Gemini 3.6 Flash Is Here: Pricing, Benchmarks and What Changed

Google released Gemini 3.6 Flash on July 21, 2026, and the headline number is not a benchmark. It is the price. Output tokens now cost $7.50 per million, down from $9.00 on Gemini 3.5 Flash, while input pricing holds steady at $1.50 per million. On top of that, the model burns 17% fewer output tokens to finish the same work, so the real saving lands well below what the sticker price alone suggests.

The launch was bigger than one model. Google shipped Gemini 3.5 Flash-Lite and a security-tuned Gemini 3.5 Flash Cyber on the same day. It also confirmed that Gemini 3.5 Pro is still stuck in partner testing, and quietly revealed that pre-training has already begun on Gemini 4. This article covers the full spec sheet, every published benchmark, what the critics are already saying, where you can actually use the model today, and whether switching is worth your time.

The Key Takeaways

  • Gemini 3.6 Flash launched July 21, 2026 at $1.50 input / $7.50 output per million tokens. Input price is unchanged; output drops from $9.00.
  • It uses 17% fewer output tokens than Gemini 3.5 Flash, up to 65% fewer on DeepSWE, and generates them at 304 per second. Only the new Flash-Lite is faster, at 350.
  • The knowledge cutoff jumps from January 2025 to March 2026, a 14-month leap that is arguably the biggest practical upgrade.
  • It beats 3.5 Flash on every published benchmark, including 58.7% vs 55.1% on SWE-Bench Pro. It also outscores Google’s own Gemini 3.1 Pro on the Artificial Analysis Intelligence Index, 50 to 46.
  • The catch: independent testing scores it at 50 on the Artificial Analysis Intelligence Index, exactly the same as Gemini 3.5 Flash. It is faster and cheaper, not smarter.

What Is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s fast, low-cost tier model, positioned for agentic coding, computer use and long-document reasoning rather than raw frontier intelligence. It sits below the Pro line in Google’s hierarchy but handles the same 1-million-token context window, which is the specification most people care about in practice.

Google’s framing in the official launch announcement leans hard on the agentic era, and the benchmark selection reflects that. The model was tuned for multi-step orchestration, full-stack refactoring and IDE agent work, not for winning academic reasoning tests.

The full spec sheet

SpecificationGemini 3.6 Flash
Model IDgemini-3.6-flash
Input context window1,048,576 tokens (1M)
Output limit65,536 tokens (64k)
Knowledge cutoffMarch 2026
Input modalitiesText, image, video, audio, PDF
Output modalitiesText
Thinking modeSupported
Function callingSupported
GroundingGoogle Search and Google Maps
Structured outputSupported
Context cachingSupported
Batch processingSupported
ReleasedJuly 21, 2026

The exact figures come from the Gemini API model documentation, which is the authoritative source if you are wiring this into production code. To see how 3.6 Flash compares with every other Gemini model, read our full Gemini model comparison.

The knowledge cutoff jump nobody is talking about

Gemini 3.5 Flash had a knowledge cutoff of January 2025. Gemini 3.6 Flash moves that to March 2026. That is a 14-month advance, and for everyday users it matters more than any percentage point on a coding benchmark.

A model frozen in January 2025 had no idea that GPT-5.6, Grok 4.5 or Claude Sonnet 5 existed. It did not know current pricing for the tools you use, recent framework releases, or anything about the last year of news. The 3.6 Flash version knows all of it, which cuts down on confidently wrong answers about anything recent.

Gemini 3.6 Flash Pricing

The pricing story is more specific than most coverage suggests. Input pricing did not move. Only output pricing was cut, and it was cut by exactly $1.50 per million tokens.

ModelInput per 1M tokensOutput per 1M tokensBest for
Gemini 3.6 Flash$1.50$7.50Agentic coding, computer use, long context
Gemini 3.5 Flash (previous)$1.50$9.00Superseded by 3.6 Flash
Gemini 3.5 Flash-Lite$0.30$2.50High-volume, latency-sensitive work
Gemini 3.5 Flash CyberNot publishedNot publishedVulnerability discovery (gated pilot)

The saving compounds in a way the table cannot show. Because the model also generates 17% fewer output tokens to complete the same tasks, you pay a lower rate on a smaller number of tokens. Independent measurement from Artificial Analysis puts the blended cost at $1.16 per million tokens against $1.31 for Gemini 3.5 Flash.

For agentic workloads the gap widens further. Google reports up to a 65% reduction in output tokens on the DeepSWE benchmark, driven by the model taking fewer reasoning steps and fewer tool calls per task. If your bill is dominated by long agent runs, that is where the money actually goes. Our Gemini pricing guide for 2026 covers how the API rates line up against the consumer subscription plans.

Gemini 3.6 Flash Benchmarks

Google published a broad benchmark set, and Gemini 3.6 Flash wins every single row against its predecessor. The margins range from modest to substantial depending on what is being measured.

Against Gemini 3.5 Flash

BenchmarkWhat it measuresGemini 3.6 FlashGemini 3.5 Flash
SWE-Bench ProReal-world software engineering58.7%55.1%
DeepSWE v1.1Agentic coding49%37%
Terminal-Bench 2.1Command-line task completion78.0%76.2%
MLE-BenchMachine learning engineering63.9%49.7%
GDPval-AA v2Economically valuable knowledge work14211349
OSWorld-VerifiedComputer use83.0%78.4%
GDM-MRCR v2 (128k)Long-context retrieval91.8%77.3%
GDM-MRCR v2 (1M)Long-context retrieval at 1M54.0%26.6%
CharXiv (no tools)Chart and figure reasoning85.2%84.2%
CharXiv (with tools)Chart reasoning with tool access89.4%84.9%

The standouts are MLE-Bench, where the jump is more than 14 points, and long-context retrieval at the full 1M window, where the score doubles. The 1M-token result is worth reading carefully, though. A 54.0% score means the model still misses roughly half of what it should retrieve at maximum context length, so treat 1M-token prompts as a capability rather than a guarantee.

Against GPT-5.6, Grok 4.5 and Claude Sonnet 5

This is where the picture gets honest. Gemini 3.6 Flash does not sweep the field against current frontier models.

BenchmarkGemini 3.6 FlashGPT-5.6 LunaGrok 4.5Claude Sonnet 5
SWE-Bench Pro58.7%Not published64.7%Not published
DeepSWE v1.149%67%Not publishedNot published
Terminal-Bench 2.178.0%84.7%Not publishedNot published
MLE-Bench63.9%Not publishedNot published66.9%
GDPval-AA v21421Not publishedNot published1607
OSWorld-Verified83.0%Not publishedNot publishedNot published

Gemini 3.6 Flash leads on OSWorld-Verified computer use, both CharXiv Reasoning variants and both long-context retrieval tests. It loses on SWE-Bench Pro to Grok 4.5, on DeepSWE and Terminal-Bench to GPT-5.6 Luna, and on both MLE-Bench and GDPval-AA v2 to Claude Sonnet 5.

That is a reasonable outcome for a cheap-tier model going up against flagships, and the price difference is large enough that the comparison is not really apples to apples. If you want the full field ranked, our best AI models list tracks where every current model stands.

Independent Testing: Where Gemini 3.6 Flash Actually Ranks

Google’s own benchmark table only compares the model to its predecessors. Independent measurement from Artificial Analysis puts it against the whole field, and the picture is more interesting.

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index chart showing Gemini 3.6 Flash scoring 50, ahead of Gemini 3.1 Pro Preview at 46
Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, ahead of Google’s own Gemini 3.1 Pro Preview at 46. Source: Artificial Analysis, 21 July 2026.

The Intelligence Index v4.1 is a composite of nine separate evaluations, including GDPval-AA v2, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt and AA-LCR. Gemini 3.6 Flash scores 50, placing it 11th among the 26 leading models Artificial Analysis charted on launch day.

RankModelIntelligence Index
1Claude Fable 5 (with fallback)60
2GPT-5.6 Sol (max)59
3Kimi K357
4Claude Opus 4.8 (max)56
5GPT-5.6 Terra (max)55
6Grok 4.5 (high)54
7Claude Sonnet 5 (max)53
8=GPT-5.6 Luna (max)51
8=GLM-5.2 (max)51
8=Muse Spark 1.1 (xhigh)51
11Gemini 3.6 Flash50
12Gemini 3.1 Pro Preview46
12=Qwen3.7 Max46
14MiniMax-M344
14=DeepSeek V4 Pro (max)44
16MiMo-V2.5-Pro42
20Gemini 3.5 Flash-Lite36
22Claude 4.5 Haiku30

The detail Google did not put in a press release is in row 12. Gemini 3.6 Flash scores 50 against 46 for Gemini 3.1 Pro Preview. Google’s cheap Flash model now outscores its own Pro-tier model on the composite index, which says as much about how long the Pro line has been stalled as it does about Flash.

The gap to the actual leaders is real but narrow. Four points separate Gemini 3.6 Flash from Claude Sonnet 5 and ten from Claude Fable 5, and those models cost considerably more per token.

Output speed

Output speed chart showing Gemini 3.5 Flash-Lite at 350 and Gemini 3.6 Flash at 304 output tokens per second
Gemini 3.5 Flash-Lite (350 t/s) and Gemini 3.6 Flash (304 t/s) hold the two fastest slots measured. Source: Artificial Analysis, 21 July 2026.

On raw generation speed the Flash family owns the top of the table. Gemini 3.5 Flash-Lite is the fastest model Artificial Analysis measured at 350 output tokens per second, with Gemini 3.6 Flash second at 304.

ModelOutput tokens per second
Gemini 3.5 Flash-Lite350
Gemini 3.6 Flash304
gpt-oss-120b (high)303
Qwen3.7 Max205
GLM-5.2 (max)201
GPT-5.6 Luna (max)190
GPT-5.6 Terra (max)135
Gemini 3.1 Pro Preview121
Claude 4.5 Haiku97
Claude Sonnet 5 (max)84
Grok 4.5 (high)69
GPT-5.6 Sol (max)63
Claude Opus 4.8 (max)59
Kimi K334

Google now holds the two fastest slots in the industry, and the spread is not close. Gemini 3.6 Flash generates tokens roughly 3.6 times faster than Claude Sonnet 5 and about five times faster than Claude Opus 4.8. For anything that streams output to a user or runs thousands of agent steps, that difference compounds into real time saved.

What the Critics Say About Gemini 3.6 Flash

The reception has not been uniformly warm, and the criticism is worth taking seriously because it comes from independent measurement rather than vibes.

Artificial Analysis scores Gemini 3.6 Flash at 50 on its Intelligence Index, ranking it 21st out of 187 models tested overall. The problem is that Gemini 3.5 Flash also scores 50. On that composite measure, the new model is not smarter than the one it replaces. Outlets including WCCFTech have run with that finding and framed the release as underwhelming.

The leaderboard backs up the complaint. Muse Spark 1.1, GLM-5.2 and GPT-5.6 Luna all score 51, Claude Sonnet 5 scores 53, Grok 4.5 scores 54 and GPT-5.6 Terra scores 55. Several of those have been available for weeks or months. Gemini 3.6 Flash arrived as a new model and landed behind all of them.

Gemini 3.6 Flash generates tokens very fast, at 303.6 tokens per second against 165.4 for 3.5 Flash. But its time to first token measures 11.54 seconds, well above the roughly 2.8-second median for reasoning models in its price bracket. It thinks for a while before it starts talking, which makes it feel slower in interactive chat than the throughput number implies.

Early community testing has also reported weak results on frontend generation and spatial reasoning prompts, though testers themselves flagged that this could reflect a lower default thinking budget rather than the model’s ceiling. Treat those reports as anecdotal until more systematic evaluation lands.

The fair summary is that Google optimised for cost, speed and token efficiency this cycle rather than for raw capability. That is a legitimate engineering goal, and for high-volume production workloads it is arguably the more useful one. It is just not the same thing as a smarter model, and the marketing does blur that line.

Where You Can Use Gemini 3.6 Flash

Google’s own surfaces

Gemini 3.6 Flash went live immediately across the Gemini app, Google AI Studio, the Gemini API, Android Studio, Google Antigravity and the Gemini Enterprise Agent Platform. Google says the new models appear in the Gemini app for everyone, but it has not published a tier-by-tier breakdown of limits, so treat free-tier quotas as unconfirmed for now. Our guide to whether Gemini is free explains how the plan structure works today.

GitHub Copilot

The model landed in GitHub Copilot on launch day for Copilot Pro, Pro+, Max, Business and Enterprise subscribers. You can select it in Visual Studio Code, Visual Studio, the Copilot CLI, the GitHub Copilot cloud agent, the Copilot app, JetBrains IDEs, Xcode and Eclipse.

One practical catch for teams. Copilot Business and Enterprise administrators have to enable the Gemini 3.6 Flash Preview policy in Copilot settings before anyone in the organisation can pick the model. If it is missing from your model list, that setting is the first place to check. Usage is billed at provider list pricing under usage-based billing.

The Other Two Models Google Shipped

Gemini 3.5 Flash-Lite

Flash-Lite is the volume play, priced at $0.30 input and $2.50 output per million tokens. At 350 output tokens per second it is the fastest model Artificial Analysis currently measures, beating every flagship on the market by a wide margin. The generational gains are dramatic too. It scores 54% on Terminal-Bench 2.1 against 31% for the previous Flash-Lite, and 1140 on GDPval-AA v2 against 642. Long-context retrieval climbs to 72.2% from 60.1%. Google is also rolling it into Google Search.

The trade-off shows up on the Intelligence Index, where Flash-Lite scores 36 against 50 for Gemini 3.6 Flash. You are buying speed and volume, not reasoning depth.

Gemini 3.5 Flash Cyber

Flash Cyber is a specialist, fine-tuned to find, validate and patch software vulnerabilities, and it runs inside Google’s CodeMender agent with multiple instances working in parallel. In Google’s published results against the V8 JavaScript engine, it surfaced 55 unique confirmed issues, compared with 47 for Gemini 3.5 Flash and 36 for Claude Opus 4.6.

You almost certainly will not get to use it. Google says the model will be available exclusively to governments and trusted partners through CodeMender, as part of a limited-access pilot programme launching soon. That is a defensible call for a model built specifically to discover exploitable bugs.

Gemini 3.5 Pro Is Late and Gemini 4 Has Started Training

The most interesting part of this launch is what was missing. Google promised Gemini 3.5 Pro at I/O in May 2026 for a June release. It has not shipped. Google now says only that the model continues to test with partners and will roll out broadly soon, with no firm date attached.

That reframes the whole announcement. Rather than competing at the top of the market, Google spent this cycle competing on price and speed in the Flash tier while its flagship slipped. Three Flash models in one day reads differently once you know the Pro model was supposed to be the story.

Buried in the same announcement was the real headline for anyone watching the long game. Google DeepMind said it has started its most ambitious pre-training run yet, for Gemini 4, and is excited by the progress. No timeline, no specifications, just confirmation that the next generation is underway. We track every credible signal in everything we know about Gemini 4.

Should You Switch to Gemini 3.6 Flash?

If you are already calling Gemini 3.5 Flash through the API, switch. There is no meaningful argument for staying. You get identical measured intelligence, better scores on every published benchmark, nearly double the output speed, a knowledge cutoff 14 months fresher, and a lower bill on both the per-token rate and the token count. The only change required is the model ID.

If you are choosing between providers, the decision is narrower. Gemini 3.6 Flash is not the most capable model available, and GPT-5.6, Grok 4.5 and Claude Sonnet 5 each beat it on specific benchmarks. It wins on computer use, long context and cost per unit of work. For high-volume agentic pipelines where you are paying for millions of tokens a day, that combination is hard to argue with. For a single hard reasoning problem, reach for a flagship instead.

The broader lesson is that no single model wins everything, and the cheap tier now moves fast enough that today’s answer expires in weeks. That is exactly why running several models side by side is more practical than committing to one. Fello AI for Mac takes that approach, putting multiple frontier models behind one $9.99/month subscription so you can move between them per task instead of paying separate bills. Our breakdown of when to use which AI model is a good starting point if you are building that habit.

For the deeper history of how this tier developed, our Gemini 3.5 Flash review covers the model 3.6 Flash just replaced.

FAQ

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s fast, low-cost AI model released on July 21, 2026. It costs $1.50 per million input tokens and $7.50 per million output tokens, handles a 1-million-token context window, and is built for agentic coding, computer use and long-document reasoning.

Is Gemini 3.6 Flash better than Gemini 3.5 Flash?

It beats Gemini 3.5 Flash on every benchmark Google published, including 58.7% against 55.1% on SWE-Bench Pro and 83.0% against 78.4% on OSWorld-Verified. It is also faster and cheaper. However, Artificial Analysis scores both models at 50 on its Intelligence Index, so the two are equally capable on that composite measure.

How much does Gemini 3.6 Flash cost?

API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Input pricing is unchanged from Gemini 3.5 Flash, while output pricing dropped from $9.00. Gemini 3.5 Flash-Lite is cheaper at $0.30 input and $2.50 output.

What is the knowledge cutoff for Gemini 3.6 Flash?

March 2026, up from January 2025 on Gemini 3.5 Flash. That 14-month jump means the model knows about events, software releases and pricing changes the previous Flash generation had no awareness of.

When is Gemini 4 coming out?

Google has not announced a release date. In the Gemini 3.6 Flash announcement, Google DeepMind confirmed it has started its most ambitious pre-training run yet for Gemini 4, but shared no timeline or specifications.

Share Now!

Facebook
X
LinkedIn
Threads
Courriel

Recevez des conseils exclusifs sur l'IA dans votre boîte de réception !

Gardez une longueur d'avance grâce à des informations sur l'IA fiables et éprouvées par les meilleurs professionnels de la technologie !