Google made Gemini 3.8 Flash generally available on September 2, 2026, three weeks after Gemini 3.7 Flash and six weeks after 3.6 Flash. It costs $0.75 per million input tokens and $3.75 per million output tokens, scores 73.7% on DeepSWE v1.1, and wins 8 of the 14 benchmark rows in Google's own comparison table.

The launch coverage settled on a single line: a cheap Flash model that beats Claude Opus 5. Google's own evaluation document tells a more useful story. Opus 5 still wins five of those rows, one of them by 32.7 points, and the DeepSWE figure republished across most of the write-ups is not the number Google published. This article has the full table and what each benchmark measures. It covers the footnote that doubles your bill on January 1, the caveats Google buried in a PDF, and whether to move off 3.7 Flash.

The Key Takeaways

  • Released September 2, 2026 as gemini-3.8-flash, Google's third Flash launch in 43 days. Gemini 3.5 Pro, promised for June, still has no entry in the API changelog.
  • It beats Gemini 3.7 Flash on every single published row. The largest jump is BioMysteryBench Human Difficult, from 43.5% to 56.5%, followed by Terminal-bench 4.0 from 11.2% to 19.1%.
  • The circulating DeepSWE number is wrong. Several write-ups report 71.0%. Google's evaluation PDF says 73.7%, against 74.0% for Claude Opus 5. That is a 0.3-point gap, not a three-point one.
  • Opus 5 still wins the hard agentic tests. Terminal-bench 4.0 goes 51.8% to 19.1%, OSWorld-2.0 75.4% to 59.0%, and GDPVal-AA v2 1824 to 1545 on Elo.
  • The price is a promotion. $0.75 and $3.75 hold only until December 31, 2026. On January 1 they become $1.50 and $7.50, and context caching doubles with them.

What Is Gemini 3.8 Flash?

Vom Herausgeber

Jedes KI-Modell in einer App

Fello AI vereint GPT-5.6, Claude 5, Gemini 3.6, Grok 4.5 und mehr in einer nativen App für Mac und iPhone.

Jetzt herunterladen!

Gemini 3.8 Flash is Google's fast, low-cost tier model, tuned this cycle for long-horizon software engineering, autonomous agents and document-heavy enterprise work. Google's API release notes describe it as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows". That entry, dated September 2, 2026, is the only hard-dated primary record of the launch, because Google shipped the model into its documentation before the marketing caught up.

The positioning is a continuation rather than a reset. Where 3.7 Flash was pitched at coding, 3.8 Flash is pitched at agents that run for a long time without supervision. That is why Google's benchmark selection leans on terminal harnesses, computer use and professional agent tasks rather than raw reasoning.

The full spec sheet

The specifications live on the model card and the API model page, not in any announcement, which is why most launch-day coverage skipped them.

SpecificationGemini 3.8 Flash
Model codegemini-3.8-flash
ReleasedSeptember 2, 2026
Input context window1,048,576 tokens (1M)
Output limit65,536 tokens (64k)
Knowledge cutoffMarch 2026, some domains only to January 2025
Input modalitiesText, image, video, audio, PDF
Output modalityText only
Thinking levelsLow, medium, high (minimal not supported)
Not supportedAudio generation, image generation, Live API
Input price$0.75 per 1M tokens (introductory)
Output price$3.75 per 1M tokens (introductory)

Every one of those figures is identical to Gemini 3.7 Flash. The context window, the output ceiling, the cutoff, the modality list and the launch price all carry over unchanged. Nothing was traded away for the benchmark gains, and nothing was added either. This release is entirely about what the model does with the same envelope.

The codename Google never confirmed

In the weeks before launch, the model circulated under the internal name Skimaki, reported by the Wall Street Journal and repeated across leak sites. The same reports claimed Google engineers preferred it to Opus 5 in internal testing. Google has never confirmed the codename or the internal comparison, and neither appears in any of its published material. Now that real numbers exist, the leaks are no longer the interesting part.

Gemini 3.8 Flash Benchmarks in Full

Google published its evaluation results as an image inside a PDF rather than as text on a web page. That is why the coverage keeps quoting three or four numbers from the marketing page instead of the table. Here is the whole thing, transcribed from Google's own Gemini 3.8 Flash model evaluation, dated September 2026. The comparison columns are Google's choice of rivals, not ours.

BenchmarkWhat it measures3.8 Flash3.7 FlashOpus 5GPT-5.6 Sol
DeepSWE v1.1Long-horizon software engineering73.7%65.3%74.0%72.7%
GDPVal-AA v2Knowledge work (Elo)1545148218241710
Vals Finance Agent v2Financial analyst tasks61.4%59.0%58.6%53.8%
Harvey's Legal AgentComplex legal workflows10.0%8.8%6.7%2.5%
Terminal-bench 2.1Agentic terminal coding89.4%85.8%89.1%88.8%
Terminal-bench 4.0General agent capabilities19.1%11.2%51.8%37.3%
GDP.PDFExpert PDF comprehension35.0%34.0%37.0%40.0%
CharXiv ReasoningSynthesis from complex charts86.2%84.5%83.7%85.8%
LVBenchLong video understanding87.8%85.4%75.4%82.1%
HLE-VerifiedMultidisciplinary expert reasoning54.9%53.6%54.4%54.5%
OSWorld-2.0Agentic computer use59.0%50.6%75.4%62.6%
BioMysteryBench (solvable)Bioinformatics research88.8%87.1%90.1%79.5%
BioMysteryBench (difficult)Bioinformatics research56.5%43.5%49.4%44.7%
LABBench2Real-world biology tasks86.2%82.1%84.2%82.1%

LVBench is quoted at 87.8% in agentic mode and 87.1% static. Google's table also carries Claude Sonnet 5 and GPT-5.6 Terra columns, both of which 3.8 Flash beats on almost everything, so they are left out here for width.

The figure most coverage got wrong

Several aggregator write-ups published on launch day report the model at 71.0% on DeepSWE v1.1. Google's evaluation PDF says 73.7%. The difference matters because the number sits next to Opus 5's 74.0%. At 71.0% the story is "Opus 5 comfortably ahead". At 73.7% it is a statistical tie between a $5 model and a $0.75 one. Read anything quoting 71.0% with that in mind. At least one write-up also labels the benchmark DeepSWE v1, where Google's methodology section is explicit that this is v1.1.

Where Opus 5 still wins

Two of the five Opus 5 wins are close enough to ignore. DeepSWE is 0.3 points, and BioMysteryBench on human-solvable problems is 1.3. The other three are not close at all. Terminal-bench 4.0, which measures general agent capability rather than coding specifically, goes 51.8% to 19.1%, a gap so wide that GPT-5.6 Sol and even GPT-5.6 Terra also finish ahead of Gemini. OSWorld-2.0 on computer use goes 75.4% to 59.0%. GDPVal-AA v2, which scores general knowledge work on an Elo scale, gives Opus 5 1824 against 1545.

The pattern is consistent. This model is excellent at bounded professional tasks with a clear finish line, in finance, law, biology, charts and video. It falls a long way behind on open-ended agent work where the model has to decide for itself what to do next. If your workload looks like the second category, the price advantage does not rescue it.

The Caveats in Google's Own Methodology

The same PDF carries a methodology section that nobody has reported, and it is unusually candid about how these numbers were produced.

On computer use, Google writes that its OSWorld-2.0 runs "were completed before the official OSWorld 2.0 team's 08.08 patch to be compatible with competitor's self-reported numbers". The Opus 5 figure there was "taken from the Fable 5.1 blog post" rather than run in-house. On long video, Google used "1024 frames for Gemini and GPT-5.6 models and 300 frames for Claude models due to API limitations". The LVBench win over Opus 5 was therefore scored with Claude seeing roughly a third of the footage. On expert reasoning, "a significant proportion of questions were blocked by content policy filters for Sonnet 5 and a small number for Opus 5", which depresses competitor scores for reasons unrelated to capability.

None of this is misconduct, and disclosing it is better practice than most vendors manage. It does mean the comparison columns are a mix of Google's own runs and rivals' self-reported figures gathered under different conditions. Treat the Gemini-versus-Gemini column as solid and the cross-vendor rows as directional. For the wider picture across every current release, our best AI models rankings track the field as a whole.

Gemini 3.8 Flash Pricing and the January Deadline

Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens. That is introductory pricing and it expires on December 31, 2026. From January 1, 2027, Google's Gemini API pricing page lists $1.50 and $7.50. Context caching doubles at the same moment, from $0.075 to $0.15 per million tokens, with cache storage going from $0.50 to $1.00 per million tokens per hour. Google's pricing table also labels the output rate as including thinking tokens, so the thinking level you choose is part of the bill, not a free extra.

Google states this plainly in the footnote of its own benchmark table, so it is not hidden, but almost every launch write-up reported the promotional rate as though it were the price. If you are modelling costs for next year, model the January number. Our full guide to Gemini pricing across plans and the API covers how the tiers interact.

How it compares on cost

Google's own table makes the cost argument better than any external comparison could, because it puts rival list prices in the same grid.

ModelInput / 1MOutput / 1M
Gemini 3.8 Flash$0.75 ($1.50 from January)$3.75 ($7.50 from January)
Claude Sonnet 5$2.00$10.00
GPT-5.6 Terra$2.00$12.00
GPT-5.6 Sol$4.00$20.00
Claude Opus 5$5.00$25.00
DeepSeek V4-Flash$0.44 peak, $0.22 off-peak$1.32 peak, $0.66 off-peak

Against Opus 5, it is roughly one-seventh of the price on both sides of the meter while tying it on DeepSWE. That is the real headline, and it survives the January increase, where it will still be a third of Opus 5's rate. It is not the cheapest option on the market though. DeepSeek V4-Flash undercuts it at every hour of the day, and DeepSeek's rates are not scheduled to double at new year.

Independent Testing and Real-World Speed

Vendor benchmarks are one input. Artificial Analysis runs its own suite, and at the high reasoning setting it scores Gemini 3.8 Flash 59 on Intelligence Index v4.1.1, ranking it 16th of 195 models. Gemini 3.7 Flash scores 56 on the same index version, so the two are directly comparable and the three-point gain is real rather than an artefact of a methodology change.

On throughput it measures 304.6 output tokens per second against 279.4 for its predecessor. Time to first token is 13.39 seconds, which the platform notes is higher than average for the price bracket. That combination is worth understanding before you switch: the model starts slower and then runs faster, which helps batch and agent workloads far more than it helps an interactive chat. These figures are rolling medians and move week to week, so check the version number and the date on any Artificial Analysis score you see quoted, including older ones on this site.

Where You Can Use Gemini 3.8 Flash

Developer access was live at announcement. The model is in the Gemini API, Google AI Studio, Google Antigravity and the Gemini Enterprise Agent Platform, and it supports the Batch API, Flex inference and Priority inference. The API free tier lists Flash models at rate caps intended for prototyping, and Google reserves the right to use free-tier prompts to improve its products. It is a way to try the model, not a way to run one.

Consumer access is where the picture is incomplete. Google's product page lists the Gemini app as a surface, and 9to5Google reports the rollout reaching Google AI Pro and Ultra subscribers along with AI Mode and Gemini in Google Sheets. At the time of writing, the Gemini Apps release notes carry no entry for 3.8 Flash at all, so the app-side rollout is running ahead of Google's own consumer documentation. If the model has not appeared in your picker yet, that is the likely reason. Anyone weighing the subscription can start with our breakdown of what the paid Gemini agent tiers actually deliver.

Three Flash Releases in Six Weeks

Gemini 3.6 Flash shipped on July 21, 3.7 Flash on August 13 and 3.8 Flash on September 2. That is three releases in 43 days, all in the cheap tier, each beating the last on essentially every published number.

Set against that, Gemini 3.5 Pro was announced in May 2026 with a promise to roll out the following month, and it still does not exist. There is no gemini-3.5-pro entry anywhere in the Gemini API changelog, no model ID and no price. Bloomberg reported in July that coding performance was the cause of the delay. Google has not given a new date and said nothing about Pro alongside this launch, which is the third consecutive Flash announcement to pass without mentioning it. Our tracking of the Gemini 3.5 Pro release date follows that story as it develops.

The reading that fits the evidence is that Flash has become Google's shipping vehicle. Each release is incremental, each is cheap, and together they have moved the cheap tier close enough to frontier performance that the flagship's absence is easier to overlook. Whether that was the plan or the fallback, it is now the pattern. For how the whole family lines up, see our comparison of every Gemini model.

Should You Switch to Gemini 3.8 Flash?

If you are on Gemini 3.7 Flash, switch today. The models cost exactly the same, share every specification, and 3.8 Flash wins every published benchmark row. There is no case for staying.

If you are choosing a cheap model for agentic coding or document-heavy professional work, it earns a test. The DeepSWE result against Opus 5 at a seventh of the cost is the strongest argument any budget model has made this year. If your work is open-ended computer use or general autonomous agency, the Terminal-bench 4.0 and OSWorld numbers say pay for Opus 5 instead. If price is the only variable, DeepSeek is cheaper and is not about to double.

The wider point is that the cheap tier now changes every three weeks and no model wins everywhere, so tying yourself to one provider means re-running this decision monthly. Fello AI puts Gemini on your Mac next to ChatGPT, Claude, Grok and DeepSeek in one native app. Testing a new release against whatever you use now takes a click instead of another API key.

Frequently Asked Questions

Is Gemini 3.8 Flash free?

Partly. It is free to try in Google AI Studio and on the Gemini API free tier, which is rate-limited for prototyping and lets Google use your prompts to improve its products. In the Gemini app it is not free: access reported so far is limited to Google AI Pro and Ultra subscribers. Production API use is billed per token.

Is Gemini 3.8 Flash better than Claude Opus 5?

On Google's own table it wins 8 of 14 rows and loses 5, including Terminal-bench 4.0 by 32.7 points and OSWorld-2.0 by 16.4. DeepSWE v1.1 is effectively a tie at 73.7% against 74.0%. Opus 5 is stronger at open-ended agent work, Gemini is stronger at bounded professional tasks, and it costs about one-seventh as much.

What is the Gemini 3.8 Flash context window?

1,048,576 tokens of input, or 1M, with a 65,536-token output limit. Both match Gemini 3.7 Flash exactly. The knowledge cutoff is March 2026, though the model card notes some domains reflect data only to January 2025, so treat recent events with care.

When does the Gemini 3.8 Flash introductory price end?

December 31, 2026. From January 1, 2027 the rate rises from $0.75 to $1.50 per million input tokens and from $3.75 to $7.50 per million output tokens. Context caching doubles at the same time, from $0.075 to $0.15 per million tokens, and cache storage from $0.50 to $1.00 per million tokens per hour.

What is the difference between Gemini 3.8 Flash and 3.7 Flash?

The specifications and the price are identical. Only the scores changed, and 3.8 Flash is ahead on every published row. The largest gains are BioMysteryBench Human Difficult, from 43.5% to 56.5%, DeepSWE v1.1 from 65.3% to 73.7%, and Terminal-bench 4.0 from 11.2% to 19.1%. Artificial Analysis scores them 59 against 56 on Intelligence Index v4.1.1.