
Google's Gemini 4 Argon, announced on 30 September, is Google's bid to rejoin OpenAI and Anthropic at the front of the AI race. Google's own figures put it first on DeepSWE v1.1, a test of long, real-world programming jobs, at 77.9% against 74.2% for Claude Opus 5.5 and 74.1% for OpenAI's GPT-6 Astra. Those are vendor numbers, not independent results. The bigger story is who gets it. For now only trusted cyber defenders in Google's Fairwind Program can use Argon, some of them with the cyber safety limits switched off, while the US government tests it under its voluntary pre-release scheme. Paying developers and Google AI Ultra subscribers come next, with no date given, and everyone else after that. When it does open, the introductory price of $2 per million input tokens and $10 per million output tokens matches OpenAI's, and the later price of $4 and $20 matches Anthropic's Opus 5.5. Anthropic launched its Mythos model the same staged way a month ago. Defenders first, everyone else later, is now the normal launch pattern for the most powerful models.
What happened
On Wednesday 30 September Google announced Gemini 4 Argon, which it calls its "next era of frontier intelligence". Frontier model means a lab's biggest and most capable AI, the kind that sets the pace for everyone else. Argon is aimed at long, multi-step professional work: writing and fixing software, financial and legal research, and finding security holes in code.
It is also a comeback attempt. Google's Gemini 3 put it near the top of the field at the end of 2025, but for most of 2026 the headlines belonged to Anthropic and OpenAI. Google cancelled a planned Gemini 3.5 Pro, and in August Demis Hassabis stepped down as head of DeepMind, Google's AI lab. Argon is the first flagship model since his successor, Koray Kavukcuoglu, took over.
What Google says it can do
Facts first, and these are Google's own facts. On DeepSWE v1.1, a test of long, realistic programming jobs, Google reports 77.9% for Argon, against 74.2% for Anthropic's Claude Opus 5.5 and 74.1% for OpenAI's GPT-6 Astra. Google also puts Argon first on Zapier's AutomationBench (51.3%), which checks whether an AI can carry out ordinary business tasks from start to finish, and on several finance and legal tests. Treat all of these as the seller's numbers until independent testers have had a go, which mostly they cannot yet.
One technical change matters more than it sounds. Argon can now write up to 1 million tokens in a single answer, up from 64,000. A token is a chunk of text, roughly three quarters of a word. That limit is what lets the model think for a very long time and produce something huge in one go, like rewriting a whole program.
Google gave examples from inside the company. Teams of Argon agents (copies of the model working through a task on their own) scanned how Google's servers use memory and freed up more than 300 terabytes without buying new hardware. Others are translating large codebases from C and C++ into Rust, a safer programming language, including an 800,000-line part of Google's Fuchsia operating system. Google says that work is still being checked by people before it ships.
"Safely releasing frontier capabilities at this level requires a phased approach" (Google)
Who actually gets it
Here is the part that makes this launch unusual. Today, Argon is open only to a set of trusted cyber defenders, the people who protect companies and public services from hackers, through Google's Fairwind Program. Some of them get a version with the cyber safety limits removed, so it can find and patch security flaws at full strength. Google says that in early use Argon found a serious flaw in healthcare software used by hospitals worldwide, one that earlier models had missed.
At the same time Google is letting the US government test the model through its voluntary pre-release programme. Next come paying developers and subscribers to Google AI Ultra, its top consumer plan. Google has given no date for that, or for anyone else. The model you can use in the Gemini app today is not Argon.
When it does open, the price is pointed. The introductory rate of $2 per million tokens in and $10 per million out is the same as OpenAI's GPT-6 Sol. The later rate of $4 and $20 is the same as Anthropic's Opus 5.5. Google is matching its rivals' prices rather than undercutting them, a signal that it sees Argon in the same tier.
Is this actually new?
No. In 2019 OpenAI held back the full version of GPT-2 for months, worried it would be used for fake news, and was widely mocked for being dramatic. Seven years later, staged release is standard. A month ago Anthropic shipped one model under two names: Fable 5.1 for everyone, with safeguards, and Mythos 5.1 only for vetted security teams. Google has now done much the same. The reason is the same in both cases. A model that is good at finding security holes so they can be fixed is, by definition, good at finding them so they can be exploited.
The everyday version
Think of a locksmith who invents a tool that opens almost any lock in seconds. The responsible move is to hand it first to the people who fit and repair locks, so they can find the weak doors and replace them, then sell it more widely once the worst doors are fixed. Google is doing that with software, and the locks are the code that runs hospitals, banks and power grids.
What it means
For the AI race, Google is credible at the front again, at least on paper. Analysts quoted by CNBC were positive but careful: one said Argon makes Google competitive, not the leader, and another said the real proof will be how it performs once businesses can use it in production.
For everyone else, the pattern is the story. The most capable models now arrive in a queue: security defenders and governments first, paying customers second, the public last. That is a sensible answer to a real risk, and it also means the gap between what the best AI can do and what ordinary users can touch is growing. Expect benchmark claims like these to stay unchecked for weeks at a time, because outsiders cannot test what they cannot get.
Curious about AI? Come build with us.
Oslo Vibe Coding runs free, beginner-friendly drop-ins where we build real things with AI. No one codes alone.