An Tran Solutions
An Tran Solutions
Back to Blog

Gemini 4 Argon Is Google's Strongest Model, but Security Teams Get It First

October 1, 20265 min readby An Tran
On this page

Most of the coverage of Gemini 4 Argon revolves around one question: has Google caught up with OpenAI and Anthropic?

I think that's the wrong question. The benchmarks will be argued over for weeks, and no one outside Google can verify them yet. What is already certain lies elsewhere: Google's most powerful model wasn't handed to developers first. It went to the people who patch vulnerabilities.

That detail changes what it means for an AI model to "launch." It also directly affects any business that runs a website.

What happened, date by date

  • September 2, 2026: Google launched the Fairwind Program, a limited access program for government agencies, national cybersecurity agencies, critical infrastructure operators (healthcare, telecoms, energy, finance), and security partners. According to Google, the program has more than 650 partners worldwide. The first model in the program was Gemini 3.8 Flash Cyber, paired with the CodeMender patching tool (Google).
  • September 30, 2026: Google announced Gemini 4 Argon, in a post under the name of Koray Kavukcuoglu, SVP at Google DeepMind. At announcement, the model was available only to defenders in Fairwind and to Google's internal teams (Google, 9to5Google).

Developers, businesses, and everyday users come later. Google says only that it will open access "as soon as possible," starting with paid API customers and Google AI Ultra subscribers. There's no specific date yet, and it isn't clear whether "Argon" is even the final commercial name (XDA Developers via daily.dev).

The numbers Google published, and why not to take them at face value yet

According to Google's table and 9to5Google:

BenchmarkGemini 4 ArgonNotes
DeepSWE v1.1 (software engineering)77.9%Claude Opus 5.5: 74.2% · GPT-6 Astra: 74.1%
CWE-bench v1 (patching security flaws)68%Tied for first
AutomationBench (end-to-end business tasks)51.3First place
LVBench (long video understanding)91.7%Highest to date

On top of that, the output limit rises to 1 million tokens, up from 64,000 in earlier Gemini generations (Google).

Impressive on paper. But three things need to be read alongside it.

One, this is Google's own table, and almost no one outside Google can run the model right now. A benchmark nobody can reproduce is a claim, not a result.

Two, the margins are fairly small. Beating Opus 5.5 by 3.7 percentage points on DeepSWE is meaningful, but it's not "leaving everyone behind." In practice, how you configure and wrap a model (prompts, tools, verification steps) often creates a bigger gap than that. I've written about this in depth in why the harness matters more than the model.

Three, even some people inside Google are skeptical. According to Bloomberg (as cited by byteiota), some Google employees feel Argon performs worse on real work than its benchmarks suggest, on certain coding tasks for example. Google disputes this. It comes from anonymous sources, so I'm noting it, not drawing a conclusion. But it's reason enough to wait for real-world testing before moving your whole stack to Argon.

The real news: the strongest model is now released in priority order

This is the part I think matters most, and it's the least discussed.

Google states its reasoning plainly: it trained Argon to be "highly capable at cybersecurity defense." The model can find, verify, and patch serious software vulnerabilities on its own (Google). A capability like that cuts both ways. Anything that can find a vulnerability to patch it can find one to exploit it.

So Google chose to give defenders first access, while also taking part in a voluntary US government process that lets authorities access models before release. According to the announcement, Google is also monitoring the model's internal activity to detect misuse and misaligned behavior (Google). The model is designed to refuse dangerous requests involving cyberattacks and CBRN weapons under Google's Frontier Safety Framework (XDA Developers via daily.dev).

Google isn't the only lab heading this way. Last week, Anthropic had Sonnet 5.5 fall back to an older model automatically on high-risk cybersecurity tasks (I broke that down in my post on Sonnet 5.5). Before that, OpenAI paused training after an agent escaped its sandbox. Three major labs, three different approaches, one message:

The familiar "launch day," when everyone gets a model at the same time, is fading away for the most powerful models. The gap between when a capability exists and when you get to use it will keep growing, and who gets it first will be the provider's call.

What does this have to do with a small business?

You won't have Argon in the next few weeks, and you probably don't need it yet. But this story touches you somewhere few people think about: your website.

Consider the speed. Gemini 3.8 Flash Cyber, the previous model in Fairwind, was already described by Google as helping defenders produce patches "in minutes instead of weeks" (Google). Argon is billed as more capable. When machines find bugs that fast, the number of disclosed vulnerabilities will rise, and with it the number of patches you need to install.

We've already seen this play out recently: WordPress patched a flaw that had sat dormant for nearly 10 years, and cPanel and Fortinet firewalls have had their own rounds of fixes. It would be no surprise if the pace picks up further.

The problem is that when defenders speed up, attackers soon catch up. Similar capabilities will end up in mainstream models, or in open-source models with no refusal mechanism at all. Fairwind's head start is temporary. What lasts is that the window between a patch being released and that flaw being exploited keeps getting shorter.

For a small or mid-sized business, the consequences are concrete:

  • "Build it and leave it" websites will get riskier. A plugin left un-updated for six months used to be a minor issue. That's about to change.
  • Someone needs to own updates. The question to put to your web team isn't "which AI do you use." It's: when a vulnerability is disclosed on Monday morning, how long until my site is patched?
  • Fewer third-party components means fewer ways in. Every plugin and every library is one more thing to monitor.

An attractive launch price, but budget on the real one

A small detail many will skim past. Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. After the introductory period, prices double to $4 and $20 (Google, 9to5Google).

For comparison, according to Yahoo Finance, Claude Opus 5.5 sits at $4 and $20, and GPT-6 Astra at $10 and $50.

In other words: Argon's standard price is exactly the same as Opus 5.5's. The 50% launch discount is there to win users, not a long-term price. If you're budgeting an AI product, use the post-promotion price, and, as I wrote in the Sonnet 5.5 post, compare cost per completed job, not just price per token.

The 1-million-token output limit is a double-edged sword. It lets you hand over an entire large code migration in one go. According to 9to5Google, Google is using Argon internally for C/C++-to-Rust migrations of more than 800,000 lines. But it also means a task that goes wrong can burn through far more money than before. Set spending limits before you switch it on.

Three things to do after this news

  1. Don't change your plans for a model you can't use yet. When Argon opens its API, test it on your own workloads before comparing. A vendor's benchmarks are only a starting point.
  2. Review your website update process now. Know which components your site runs, who updates them, and how fast they're updated when a new vulnerability lands. If there's no clear answer, make it a priority this quarter.
  3. Budget AI on post-promotion prices, and always set spending caps on long-running tasks.

The question "will Google win" will be answered in the coming months. The question "can my website keep up with the new pace of patching" should be answered now.

If you'd like someone to audit your website and take on regular maintenance and security updates, take a look at my services or get in touch directly.

Sources

Related articles

An Tran Solutions
September 30, 20268 min read

Spec-Driven Development Isn't Documentation. It's a Governance System for AI-Written Code

From vibe coding to spec-driven development: why longer specs won't save you, and a four-layer governance framework (constitution, risk-tiered specs, executable acceptance, control gates) that keeps AI-written code under your control. With data from Veracode, Thoughtworks and Martin Fowler.