Thumbnail for “Gemini 3.5 Review: What Google Launched at I/O 2026” showing bold amber and white headline text reading “Gemini 3.5 Review: Flash Beats Pro?” beside a glowing Gemini 3.5 app-style card on a dark blue and purple neon background, with an “I/O 2026” badge in the corner.

Gemini 3.5 Pro Release Date: What We Know (+ Flash Review)

Google split Gemini 3.5 into two launches, and only one has actually arrived. Gemini 3.5 Flash went live on May 19, 2026 at Google I/O 2026 and did something Google had never done before, with the cheap, fast “Flash” tier beating the previous flagship on coding. The other half, Gemini 3.5 Pro, is the model everyone is waiting on, and its release date keeps moving. Google’s rumored July 17, 2026 target passed with no public launch, and Pro still sits in a limited Vertex AI preview, with an August window now rumored.

This article covers both sides. You get what Gemini 3.5 Flash delivers today, the full benchmarks against Gemini 3.1 Pro, and pricing, plus everything known about Gemini 3.5 Pro, its 2-million-token context window, built-in Deep Think reasoning mode, and every clue about when it will finally ship. Google has also confirmed a custom Gemini model will power Apple’s rebuilt Siri, which raises the stakes further.

The Key Takeaways

  • Gemini 3.5 Flash launched May 19, 2026 at Google I/O 2026, the first model in the Gemini 3.5 family.
  • It beats the bigger Gemini 3.1 Pro on coding and agentic benchmarks while running 4x faster in output tokens per second.
  • API pricing is $1.50 per million input tokens and $9.00 per million output tokens, about 25% cheaper than Gemini 3.1 Pro.
  • It is free in the Gemini app and AI Mode in Search, with a 1M-token context window.
  • Gemini 3.5 Pro missed its rumored July 17, 2026 launch and remains in limited Vertex AI preview, with an August window now rumored; it brings a 2M-token context window and Deep Think reasoning.

What Is Gemini 3.5 Flash?

Gemini 3.5 Flash is Google’s new speed-and-cost model, and it is the headline release of Google I/O 2026. In its official announcement, Google calls the family “frontier intelligence with action,” and the framing is deliberate. Flash models used to be the budget option you reached for when you did not need the smart model. Gemini 3.5 Flash flips that, delivering performance Google says “rivals large flagship models on multiple dimensions” while keeping Flash-tier speed and price. Alibaba shipped its own frontier reasoner the next day; see our Qwen3.7-Max review for the head-to-head with Flash on pricing and benchmarks.

The model is tuned for coding and agentic tasks, the long-horizon work where an AI plans, calls tools, and iterates instead of answering a single question. It also leads on multimodal understanding, scoring 84.2% on CharXiv Reasoning. The API model ID is gemini-3.5-flash, the context window is 1,048,576 input tokens with 64K output tokens, and the knowledge cutoff is January 2026. If you want the background on how Google got here, our breakdown of the jump from Gemini 2.5 to Gemini 3 traces the version history.

Gemini 3.5 Flash vs Gemini 3.1 Pro: The Benchmarks

Here is the part that matters. A Flash model is not supposed to beat last generation’s Pro model, and Gemini 3.5 Flash does exactly that on the benchmarks Google cares about most. The table below compares it against Gemini 3.1 Pro, the model it effectively replaces for most users.

BenchmarkGemini 3.5 FlashGemini 3.1 ProWinnerWhat it measures
Terminal-Bench 2.176.2%70.3%3.5 FlashReal coding in a terminal
GDPval-AA1,656 Elo1,314 Elo3.5 FlashReal-world agentic tasks
MCP Atlas83.6%78.2%3.5 FlashScaled tool use
CharXiv Reasoning84.2%N/A3.5 FlashMultimodal understanding
Humanity’s Last Exam40.2%44.4%3.1 ProHardest expert reasoning
ARC-AGI-272.1%77.1%3.1 ProAbstract reasoning
Output speed~4x fasterbaseline3.5 FlashOutput tokens per second

On GDPval-AA, the jump is dramatic. Gemini 3.1 Pro scores 1,314 Elo; Gemini 3.5 Flash scores 1,656 Elo, a 342-point swing on real-world agentic work. Google says Gemini 3.5 Flash is 4 times faster than other frontier models when measured in output tokens per second. For agents that loop through dozens of steps, that speed gap compounds fast.

Where Gemini 3.1 Pro Still Wins

Gemini 3.5 Flash is not a clean sweep, and Google is honest about it. On the hardest pure-reasoning tests, the older Gemini 3.1 Pro still has the edge. It leads Humanity’s Last Exam (44.4% vs 40.2%) and ARC-AGI-2 (77.1% vs 72.1%), and it holds up better on very long-context retrieval, scoring 84.9% on MRCR v2 at 128k tokens against 77.3% for Flash, per independent benchmark and pricing data.

The practical read is simple. For coding, tool use, and agent workflows, Gemini 3.5 Flash is the better and cheaper pick today. For the deepest reasoning problems or huge-context document work, Gemini 3.1 Pro is still worth the extra cost until Gemini 3.5 Pro lands. Our full Gemini 3.1 Pro review covers where that model still shines.

Gemini 3.5 Flash Pricing

For most people, Gemini 3.5 Flash is free. It is the default model in the Gemini app and in AI Mode in Google Search at no cost. The pricing below matters if you are a developer calling it through the API. If you just want the short answer, see whether Gemini is free.

API token pricing

Token typeGemini 3.5 FlashGemini 3.1 Pro
Input (per 1M)$1.50$2.00
Output (per 1M)$9.00$12.00
Cached input (per 1M)$0.15N/A

That makes Gemini 3.5 Flash about 25% cheaper on both input and output than Gemini 3.1 Pro at Google’s official API rates ($2.00 in / $12.00 out for 3.1 Pro), while scoring higher on the coding and agentic benchmarks. It is roughly 3x the price of the old Gemini 3 Flash ($0.50 in / $3.00 out), but far more capable, so it sits in a new sweet spot. Non-global regions are priced slightly higher at $1.65 input / $9.90 output. For the wider context, see our full Gemini pricing breakdown and our AI pricing comparison across the major models.

How to Use Gemini 3.5 Right Now

Gemini 3.5 Flash is already live everywhere Google ships Gemini. You do not have to wait or join a list. There are four ways to use it today.

  1. Gemini app: open the Gemini app on web, Android, or iOS (the Gemini iOS app now ranks among the best AI apps for iPhone). Gemini 3.5 Flash is the new default model, so you are already using it.
  2. AI Mode in Search: Gemini 3.5 Flash now powers AI Mode in Google Search worldwide, free for everyone.
  3. Gemini API: developers can call gemini-3.5-flash through Google AI Studio, the Gemini API, and Android Studio.
  4. Enterprise: it is available through Google Antigravity, the Gemini Enterprise Agent Platform, and Gemini Enterprise.

If you mostly use Gemini on a Mac, our Gemini Mac app review walks through the desktop experience. Prefer one subscription that gives you several models instead of juggling apps? The Fello AI app for Mac puts Gemini, ChatGPT, Claude, Grok and DeepSeek behind a single $9.99/month plan, so you can switch models per task without paying for each one separately.

Gemini Spark and the Rest of Google I/O 2026

Gemini 3.5 Flash powers more than chat. Google used it to launch Gemini Spark, a 24/7 personal AI agent that works in the background to handle recurring tasks across your digital life. Spark rolled out to trusted testers first, followed by a Beta for Google AI Ultra subscribers in the US, as reported by Engadget. The fact that Spark runs on Gemini 3.5 Flash rather than the heavier Pro tier is what lets it stay online 24/7 at agent-tier pricing.

The keynote also brought a Neural Expressive redesign of the Gemini app, a personalized Daily Brief morning digest for AI Plus, Pro, and Ultra subscribers, and updates to Gemini Omni. Google AI Ultra now starts at $99.99/month, with a higher-limit $200 plan on top. On the safety side, Google says Gemini 3.5 ships with strengthened cyber and CBRN safeguards, fewer incorrect refusals on safe prompts, and new interpretability tooling that inspects the model’s internal reasoning before a response is returned. Google also previewed Android XR smart glasses with Gemini 3.5 Flash built in, putting Google directly against Apple in the 2027 race.

When Is Gemini 3.5 Pro Coming Out?

Gemini 3.5 Pro has not launched yet. Google announced it at I/O on May 19, targeted June for general availability, then slipped that into July after what it described as quality refinements from early enterprise testing. Prediction markets and reporting zeroed in on a July 17, 2026 date, but that day came and went with no public release. As of mid-July, Pro still sits in a limited Vertex AI enterprise preview, Google has not committed to an exact day, and reporting now points to an August window. Pichai told the I/O audience, “I know you can’t wait to get your hands on it,” but the wait has run well past what was promised.

Two things make Pro worth the wait. It ships with a 2-million-token context window, the largest of any production frontier model and double Claude Opus 4.8, plus a built-in Deep Think mode that runs extended, deliberate reasoning for the hardest problems. Pricing is not official yet. Early estimates range from around $2 per million input and $12 per million output, matching Gemini 3.1 Pro, up to as high as $15 in and $60 out, so treat any figure as a placeholder until Google publishes rates.

FeatureGemini 3.5 FlashGemini 3.5 Pro
StatusLive since May 19, 2026Delayed past July 17; Vertex preview only
Context window1M tokens2M tokens
Deep Think modeNoYes
Best forSpeed, coding, agents, everyday chatHardest reasoning, huge-context work
API price (per 1M)$1.50 in / $9.00 outNot confirmed (est. $2 to $15 in)

Until Pro is live for everyone, Gemini 3.5 Flash covers almost everything and Gemini 3.1 Pro remains the pick for the deepest reasoning tasks. We will update this article the moment Gemini 3.5 Pro goes generally available.

Gemini Is About to Power Apple’s New Siri

Gemini’s reach now extends well beyond Google’s own apps. Under a multi-year deal reportedly worth around $1 billion a year, a custom Gemini model, said to run roughly 1.2 trillion parameters inside Apple’s own data centers, powers Apple’s rebuilt Siri, the biggest overhaul the assistant has ever had. Apple unveiled the new Siri AI at WWDC 2026 as a standalone, chatbot-style app for iPhone, iPad, and Mac that can pull context from your messages, emails, and photos and run multi-step actions across apps.

The iOS 27 public beta went live in mid-July 2026, so anyone can now test the Gemini-powered Siri AI ahead of a general release in the autumn alongside the next iPhone. That means the same model family reviewed here already sits behind the assistant for early testers, and soon hundreds of millions more Apple devices. For the full picture of what is changing, see our breakdown of whether Siri is now real AI.

How Gemini 3.5 Pro and Flash Stack Up Against ChatGPT and Claude

Against the wider field, the picture is competitive rather than dominant. As of July 2026, GPT-5.6 leads on benchmarks like Terminal-Bench, while Claude Opus 4.8 and Claude Fable 5 lead the agentic coding tests like SWE-bench Verified. Gemini’s edge is speed and price at near-flagship quality. Gemini 3.5 Flash does not claim the overall crown, but it changes the value equation by delivering this tier of performance at Flash cost and speed, and Gemini 3.5 Pro is built to challenge the top tier on the hardest reasoning tasks.

For a head-to-head on everyday use, see how Gemini compares to ChatGPT, or our three-way ChatGPT vs Claude vs Gemini breakdown for Mac users. The honest takeaway is that no single model wins everything, which is exactly why model choice per task beats model loyalty.

Should You Use Gemini 3.5?

Yes, and you probably already are. If you use the Gemini app or AI Mode in Search, Gemini 3.5 Flash is your default as of today, at no cost, with a 1M-token context window. For developers, it is the obvious move for coding and agent workloads given it beats Gemini 3.1 Pro on those benchmarks at 25% lower cost and 4x the speed. The one reason to hold off is the deepest reasoning work, where holding out for Gemini 3.5 Pro is the smarter call. Open the Gemini app, ask it something hard, and watch how fast it answers.

FAQ

When did Gemini 3.5 release?

Gemini 3.5 Flash launched on May 19, 2026 at Google I/O 2026. It is the first model in the Gemini 3.5 family and is already the default in the Gemini app and AI Mode in Search.

Is Gemini 3.5 free?

Yes. Gemini 3.5 Flash is free in the Gemini app and in AI Mode in Google Search. Developers pay $1.50 per million input tokens and $9.00 per million output tokens through the API.

Is Gemini 3.5 Flash better than Gemini 3.1 Pro?

On coding and agentic benchmarks, yes. It beats Gemini 3.1 Pro on Terminal-Bench 2.1, MCP Atlas, and GDPval-AA while running about 4x faster. Gemini 3.1 Pro still leads on the hardest pure-reasoning tests.

When is Gemini 3.5 Pro coming out?

Not yet. Google announced Gemini 3.5 Pro at I/O on May 19, 2026 and aimed for June, then July. A rumored July 17 launch passed with no public release, so as of mid-July it remains in limited Vertex AI preview, with an August window now rumored. It brings a 2M-token context window and a Deep Think reasoning mode, and Google has not confirmed an exact date.

What is Gemini Spark?

Gemini Spark is a 24/7 personal AI agent powered by Gemini 3.5 Flash. It handles recurring tasks in the background and rolled out to trusted testers first, with a Beta for Google AI Ultra subscribers in the US.

Share Now!

Facebook
X
LinkedIn
Threads
Email

Get Exclusive AI Tips to Your Inbox!

Stay ahead with expert AI insights trusted by top tech professionals!