Google announced Gemini 4 Argon on September 30, 2026, the first model of the Gemini 4 generation and Google's first flagship since Gemini 3.1 Pro in February. Google calls it "our next era of frontier intelligence" and pitches it at long, complex work: software engineering, legal and financial research, and cybersecurity defense. It can write up to 1 million output tokens in a single response, up from 64,000 on earlier Gemini models, and it launches at an introductory price of $2 per million input tokens and $10 per million output tokens.

Access is limited. On launch day Gemini 4 Argon went only to a vetted group of cyber defenders in Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers are next, and Google has not given a date. This guide covers what Argon is and what Google's benchmarks show. It also examines independent results from the first hours, reported doubts among Google employees, and when you can expect to use it.

The Key Takeaways

  • A real frontier model again: Artificial Analysis scores Gemini 4 Argon 53 on its Intelligence Index, level with GPT-6 Astra and Claude Fable 5.1, and behind Claude Opus 5.5 and Sonnet 5.5. Gemini 3.1 Pro scored 30.
  • Cheap for now: $2 input and $10 output per million tokens is a 50% launch discount. The standard price is $4 and $20, the same as Opus 5.5, and Google has not said when the discount ends.
  • Where it leads: 77.9% on DeepSWE v1.1, #1 on the Vals Index and the LMArena Text Arena, and the lowest hallucination rate Artificial Analysis has measured among leading models.
  • Where it trails: Claude Opus 5.5 beats it on Terminal-Bench 4.0 (66.4% vs 57.4%) and GPT-6 Astra on FrontierSWE v2 (65.5% vs 55.0%). Bloomberg reports some Google staff think it struggles with real coding work.
  • Who gets it: only Fairwind Program cyber defenders today. Paid API users and Google AI Ultra subscribers come next, with no date. There is no public API model ID yet.

What Gemini 4 Argon Is

Do editor

Todos os modelos de IA numa só aplicação

Fello AI reúne GPT-6, Claude 5.5, Gemini 3.8, Grok 4.7 e mais numa só aplicação nativa para Mac e iPhone.

Descarregue já!

Gemini 4 Argon is Google DeepMind's new top model and the first release of Gemini 4. The launch post, written by Koray Kavukcuoglu, SVP of Google DeepMind and Google's Chief AI Architect, says Argon is "built to sustain deep reasoning across complex, long-horizon workflows." It is a reasoning model, which means it thinks through a problem before it answers, and Google names four areas where it wants Argon to be at the frontier: coding, enterprise knowledge work, cybersecurity defense and creative writing.

Google DeepMind announced it on its official account:

Why Gemini 4 took so long

Gemini 4 Argon ends a long gap in Google's lineup. Google's last flagship was Gemini 3.1 Pro, which we covered in our Gemini 3.1 Pro launch guide in February. Since then Google shipped only faster Flash models, most recently Gemini 3.8 Flash, and has still not released the Gemini 3.5 Pro it said in July was testing with partners. Kavukcuoglu told reporters on September 24 that Google "took a little bit of a step back" to focus on Flash and on Gemini 4, and that the plan was to release "an early post-training output" as soon as possible, 9to5Google reported. Our Gemini 4 tracker followed those developments from the first pre-training mention in July.

The name is new too. Earlier Gemini models were split into Pro, Flash and Flash-Lite tiers. Argon is the first Gemini with a codename-style name, like OpenAI's GPT-6 models Astra and Sol. Google has not said what the other Gemini 4 models will be called or when they will ship.

The specs that matter

These are the figures Google has published so far, plus the ones Artificial Analysis confirmed in its own testing. Google has not yet published a model card, a knowledge cutoff or an API model ID.

SpecGemini 4 Argon
AnnouncedSeptember 30, 2026
Available toFairwind Program partners only (for now)
Introductory price per 1M tokens$2 input, $10 output, $0.10 cached input
Standard price per 1M tokens$4 input, $20 output
Max output1M tokens (up from 64K)
Context window1M tokens (per Artificial Analysis)
Input / outputText and image in, text out (Artificial Analysis's launch post also lists video and speech input)
Reasoning setting testedHigh, the highest available
API model IDNot published yet

Who Can Use Gemini 4 Argon Today

Almost nobody outside Google can use Gemini 4 Argon yet. Google DeepMind is rolling it out first through the Fairwind Program, a controlled-access scheme for organizations that defend networks and software. Its page gives priority to three groups: governments and national cyber authorities, critical-infrastructure operators in healthcare, telecoms, energy and finance, and core technology platforms. Academic labs that focus on defensive benchmarking can also apply. Applicants are vetted, need phishing-resistant multi-factor authentication, and may not share or resell access. Google says more than 650 partners take part.

Google's launch post says Fairwind partners and internal Google teams get a version of Argon without its cyber guardrails, the filters that normally block requests that look like hacking. They can also use it inside CodeMender, Google's agent that finds and patches security bugs. The only partner Google names in its launch post is Wiz, the cloud security company Google owns. Wiz used Argon in its Scan for Good project and, according to Google, found a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models had missed.

Google is also taking part in the U.S. government's voluntary process for reviewing frontier models before release. Anthropic and OpenAI also limited initial access to their strongest cyber models through Anthropic's Project Glasswing and OpenAI's Daybreak program.

When everyone else gets it

Google says paid API customers and Google AI Ultra subscribers will get Gemini 4 Argon next, followed by wider access for developers, enterprises and consumers "as soon as possible." It has given no date. According to The Next Web, Kavukcuoglu said the week before launch that Google wanted Gemini 4 out well before the end of the year. Nothing Google has published so far mentions the free Gemini app or the Google AI Pro plan, so expect Argon to reach the top tier first. Our Gemini pricing guide covers what each plan costs today.

Logan Kilpatrick, who leads Google AI Studio and the Gemini API, framed the rollout the same way:

Gemini 4 Argon Pricing: Half Price for Now

Gemini 4 Argon's launch price is a 50% promotion. Google's post says Argon "will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price." After the promotion it costs $4 and $20. Google has not said how long the introductory period lasts, so budget on the standard price if you are planning production work.

At the launch price, Argon is one of the cheapest frontier models. At the standard price, it costs the same per token as Claude Opus 5.5.

ModelInput per 1MOutput per 1MCached input per 1M
Gemini 4 Argon (introductory)$2$10$0.10
Gemini 4 Argon (standard)$4$20$0.20
Gemini 3.1 Pro$2$12not compared
Claude Opus 5.5$4$20$0.20
Claude Sonnet 5.5$2$10$0.20
GPT-6 Astra$10$50$1
GPT-6.1 Sol$2$10$0.10

What a task actually costs

Models use very different numbers of tokens for the same job, so token prices alone do not tell you what a task costs. Artificial Analysis measured what it costs to run each model through its Intelligence Index. At the introductory price, Gemini 4 Argon costs $1.99 per task, 60% of GPT-6 Astra's $3.26. That is well under Claude Opus 5.5 at $5.98 and Claude Sonnet 5.5 at max effort at $7.62. At the standard price, Argon's cost rises to $3.98 per task, about 1.2 times GPT-6 Astra.

Argon is not a frugal model. Artificial Analysis says its low cost comes from low token prices, "rather than reduced token use." Argon averaged 62,000 output tokens per task, against 27,000 for GPT-6 Astra, and it costs 2.7 times as much per task as GPT-6.1 Sol. Without the discount, Argon loses most of its price advantage. Our guide to what AI tokens are explains how token counts turn into a bill.

Gemini 4 Argon Benchmarks: Google's Numbers

Google's launch post shows Gemini 4 Argon leading on long software tasks, business automation, legal and financial work, and long video. Google's full benchmark table is below, followed by the rows that matter most, checked against the write-ups by VentureBeat and MarkTechPost. By VentureBeat's count, Argon leads 12 of the 18 benchmarks Google published and ties one more.

Google's Gemini 4 Argon benchmark table comparing Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across 18 benchmarks in knowledge work, agentic coding, science, long context, computer use, multimodal understanding and cybersecurity
Google's full Gemini 4 Argon benchmark table, with blue marking Argon's leads and grey the rows a rival wins. [Source: Google]
BenchmarkGemini 4 ArgonClaude Opus 5.5GPT-6 Astra
DeepSWE v1.1 (long software tasks)77.9%74.2%74.1%
Vals Index (finance, law, coding, tax)68.9%67.0%63.1%
AutomationBench (business workflows)51.3%42.5%41.4%
Vals Finance Agent v265.4%58.6%53.5%
Harvey Legal Agent19.6%3.8%5.4%
GraphWalks (long context)84.2%66.8%71.8%
LVBench (long video)91.7%83.7%87.5%
CWE-bench v1 (fixing security flaws)68%67%68%
Gray Swan IPI (prompt-injection success, lower is better)0.7%1.0%8.5%
Terminal-Bench 4.0 (agentic coding)57.4%66.4%58.2%
FrontierSWE v255.0%62.3%65.5%
PostTrainBench45.3%49.3%44.3%
Terminal-Bench Science 0.157.6%63.3%68.1%

Argon's biggest wins are in professional knowledge work rather than pure coding: legal drafting, financial research and business automation. Its coding results are mixed. Argon leads DeepSWE, which tests long real-world software tasks. Claude Opus 5.5 is well ahead on Terminal-Bench 4.0, and GPT-6 Astra edges Argon there too. Astra leads FrontierSWE v2 by ten points. Google's post does not claim a clean sweep, and it should not.

Treat rival scores in a vendor's table with care. Labs run each other's models with their own settings, so the numbers do not always match. Anthropic's own launch table, for example, puts Claude Opus 5.5 at 40.0% on AutomationBench, not the 42.5% Google lists.

The cyber numbers

Cybersecurity is the headline use case, but Argon does not dominate the tests. On CWE-bench v1, which measures how well a model fixes known classes of security flaws, Argon scores 68%. That ties GPT-6 Astra and Grok 4.7 and sits one point above Claude Opus 5.5. The Fairwind page adds a score of 85.8% on a real-world vulnerability discovery benchmark that has not been published, so it cannot be compared with other models. On independent tests, Vals ranks Argon #2 on its CyberBench at 77.86%, 0.12 points behind GPT-6 Sol.

The prompt-injection number matters more for everyday users. On Gray Swan's indirect prompt injection benchmark, attacks hidden inside web pages or documents succeeded against Argon only 0.7% of the time, against 8.5% for GPT-6 Astra. These attacks hit AI agents browsing the web or reading your email, so a low success rate is a safety gain.

Google's Gray Swan indirect prompt injection chart: attack success rate of 0.7% for Gemini 4 Argon, 1.0% for Claude Opus 5.5 and Fable 5.1, 8.5% for GPT-6 Astra and up to 52.7% for Kimi K3
Gray Swan indirect prompt injection results, where a lower attack success rate is better. [Source: Google]

What Independent Testers Found

Three independent leaderboards published Gemini 4 Argon results within hours of the launch, even though Argon is not publicly available. All three put it at or near the top and found the same weaknesses as Google's own table.

Artificial Analysis: level with GPT-6 Astra

With reasoning set to high, Gemini 4 Argon scored 53 on Artificial Analysis's Intelligence Index. That matches GPT-6 Astra and Claude Fable 5.1 and is one point above GPT-6.1 Sol, but below Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56. It is also 23 points above Gemini 3.1 Pro Preview, which puts Google back among the top three labs on this measure:

The most useful finding is about hallucinations. On Artificial Analysis's AA-Omniscience test, Gemini 4 Argon made things up 15% of the time, the lowest rate of any model scoring 45 or more on the Intelligence Index. GPT-6 Astra scored 51% and GPT-6.1 Sol 54%. In practice, Argon is much more likely to say it does not know than to guess. The trade-off is lower accuracy: Argon answered 50% of the factual questions correctly, 13 points below GPT-6 Astra. A model that declines more often gets fewer answers wrong, but it also gets fewer right.

Artificial Analysis also found Argon much stronger at agent work than past Gemini models. It ranked #1 on AutomationBench-AA at 77.5%, six points ahead of Claude Sonnet 5.5. On Terminal-Bench 4 it scored 57%, a 53-point jump over Gemini 3.1 Pro Preview but still behind Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Astra.

Vals and LMArena: first place

On the Vals Index, which tests finance, legal, medical and coding tasks, Gemini 4 Argon ranks #1 of 41 models at 68.90%. It is ahead of Claude Sonnet 5.5 at 67.04% and Claude Opus 5.5 at 66.97%, and it cost $15.68 per test run against $32.14 for Opus 5.5. Vals also recorded a perfect 100% on its IOI programming-olympiad test, tied with GPT-6 Astra.

In LMArena's blind Text Arena, where people vote between two anonymous answers, Argon took first place with 1525 points. It also ranked first in the coding, hard prompts, instruction following and creative writing categories. Its web development score was weaker: #8 in Code Arena: WebDev with 1679 points, although that is a big jump from Gemini 3.8 Flash at #29.

If you want the background on what these leaderboards measure and where they mislead, see our AI benchmarks explainer.

The Doubts: Great on Benchmarks, Weaker at Real Work?

On launch day, Bloomberg reported that some Google employees are skeptical of Gemini 4. Bloomberg cited people with direct access who said the model does well on industry benchmarks but less well on employees' real tasks. They also said it struggles with some coding jobs. One weakness they named was front-end design, the code that decides how apps and websites look. Google disputed the report and said it would be inaccurate to call Gemini 4 an underperformer in areas like coding. One employee told Bloomberg there is a "large consensus" inside the company that the model is at the frontier.

The public data partly supports both sides. Argon's #8 finish in Code Arena: WebDev matches the front-end complaint, and its Terminal-Bench and FrontierSWE scores trail Anthropic and OpenAI. But its DeepSWE lead and first place in the Text Arena's coding category show it is not a weak coder overall. Until paying users can try it, nobody outside Google and Fairwind can say how the benchmark lead holds up in daily work.

How Google Is Already Using Argon

Google says teams across the company are already using Gemini 4 Argon, and its launch post gives four examples. Sundar Pichai, Google's CEO, said teams use it "extensively at Google, from coding to quantum computing" with "great feedback":

  • Rust migrations: Argon agents are moving C and C++ codebases to the memory-safe Rust language, including work on the Zircon kernel of Google's Fuchsia operating system, which MarkTechPost puts at more than 800,000 lines.
  • Faster video decoding: in the libgav1 AV1 decoder, agents replaced about 32,000 lines of hand-tuned SIMD code and made the decoder 2.7 times faster.
  • Data-center memory: Argon analyzed profiling data from Google's server fleet and applied memory optimizations on its own, freeing more than 300 TiB.
  • Quantum computing: Argon beat a published quantum algorithm baseline by 40% in an optimization task.

These examples come from Google and cannot be independently checked. They show what the 1 million token output limit is for: jobs that run for hours and write huge amounts of code in one go.

What the 1 million token output limit means

Most models can write between 64,000 and 128,000 tokens in one response. Gemini 4 Argon can write up to 1 million, roughly 750,000 words. For chat this makes no difference. For agents that migrate a whole codebase, draft a long legal document or reason through a hard problem for many minutes, it removes a limit that used to cut work off halfway. Artificial Analysis says the Gemini API has a new feature called Long Decode Continuation. It pauses very long answers and resumes them across follow-up calls, letting reasoning reach 1 million tokens without the request timing out.

Long outputs also mean long waits and big bills. Every output token is billed at the $10 or $20 output rate, so a single million-token answer costs $10 at the launch price and $20 after it.

Safety: Guardrails On for Everyone but Defenders

Google says it has strengthened four kinds of safeguards before Gemini 4 Argon's wider release. They cover misuse for cyberattacks and chemical, biological, radiological and nuclear weapons, attacks hidden in content the model reads, signs that the model is working against its instructions, and insecure agent environments. For the last two, Google says it monitors the model's chain of thought and runs high-risk tests in sealed sandboxes. The general release will keep the cyber guardrails that Fairwind partners and internal Google teams work without.

Google has not yet published a full safety report or model card for Argon. Anthropic published a system card for Claude Opus 5.5 on launch day, which our Opus 5.5 guide summarizes. When Google publishes one for Argon, it will be the best source for how the model behaved in Google's own red-team tests.

Gemini 4 Argon vs Claude Opus 5.5 vs GPT-6 Astra

You cannot pick Gemini 4 Argon for real work yet, but the early data already shows where it will fit. This is how the three frontier models compare on the evidence so far.

If you needBest pick on current dataWhy
Legal, finance and business automationGemini 4 ArgonClear leads on Harvey Legal Agent, Finance Agent v2, AutomationBench and the Vals Index
Answers you can trust not to be inventedGemini 4 Argon15% hallucination rate, the lowest Artificial Analysis has measured among leading models
Long video and very long documentsGemini 4 ArgonLeads LVBench and GraphWalks
Terminal and agentic codingClaude Opus 5.566.4% on Terminal-Bench 4.0 against Argon's 57.4%
Front-end and web designClaude Opus 5.5 or GPT-6 AstraArgon is #8 in Code Arena: WebDev
Hard science and frontier codingGPT-6 AstraLeads FrontierSWE v2 and Terminal-Bench Science
Highest overall intelligence scoreClaude Opus 5.558 on the Artificial Analysis Intelligence Index against 53
Lowest cost per taskGPT-6.1 SolArgon costs 2.7 times as much per task even at its launch price

If you want to compare models on your own prompts, you can already run the current Gemini, GPT and Claude models side by side in Fello AI and switch between them in one chat.

The Verdict on Gemini 4 Argon

Gemini 4 Argon puts Google back at the frontier after seven months without a flagship. It leads on legal, financial and automation work, has the lowest hallucination rate among top models, and at the launch price it gets GPT-6 Astra-level results for 60% of the cost per task. It does not beat Claude Opus 5.5 overall, and it trails on terminal coding and front-end work, the same weak spot Google's own employees reportedly flagged. For most people, access is the bigger issue: you cannot use it yet. Wait for the API and AI Ultra rollout, check whether the discount still applies, and test Argon on your own work before you switch.

Frequently Asked Questions

When was Gemini 4 Argon released?

Google announced Gemini 4 Argon on September 30, 2026. On that day it went only to trusted cyber defenders in Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers will get it next, but Google has not given a date.

How much does Gemini 4 Argon cost?

Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens during an introductory period, with cached input at $0.10. After that it costs $4 and $20, the same as Claude Opus 5.5. Google has not said when the introductory price ends.

Can I use Gemini 4 Argon in the Gemini app?

Not yet. At launch, Gemini 4 Argon is available only through the Fairwind Program, and Google has not mentioned the Gemini app. Google says Google AI Ultra subscribers will be among the first to get it after the Fairwind rollout. The company has not said when Argon will reach the Pro or free plans.

Is Gemini 4 Argon better than Claude Opus 5.5?

Gemini 4 Argon beats Claude Opus 5.5 on legal, finance, automation and long-video benchmarks and on the Vals Index. Opus 5.5 is ahead on Terminal-Bench 4.0, FrontierSWE v2 and the Artificial Analysis Intelligence Index, where it scores 58 against Argon's 53. Which is better depends on the kind of work you do.

What is the Fairwind Program?

The Fairwind Program is Google DeepMind's early-access scheme for vetted cyber defenders, such as governments, critical-infrastructure operators, technology platforms and academic security labs. Members get Gemini 4 Argon without its cyber guardrails for defensive work. Google says it has more than 650 partners.

What happened to Gemini 3.5 Pro?

Google has not released Gemini 3.5 Pro and has not said whether it will. Google DeepMind's Koray Kavukcuoglu said in September 2026 that the company "took a little bit of a step back" to focus on its Flash models, and that its focus now is Gemini 4. Gemini 4 Argon is the first Google flagship since Gemini 3.1 Pro in February 2026.