Google unveils Gemini 4 Argon, but the new flagship is not open yet
Gemini 4 Argon
Google’s new AI model for difficult, long-running work.
benchmark
A fixed test used to compare AI systems.
Fairwind Program
Google’s early testing program for trusted cybersecurity defenders.
What happened
On September 30, Google introduced Gemini 4 Argon. Google DeepMind, Google’s AI research division, developed the model. Google calls Argon its most advanced AI model yet.
The model targets real software engineering, legal and finance work, and cybersecurity defense. Its main promise is handling long, complicated workflows. It is meant to keep track of many steps, rather than answer one question and stop.
Argon is not open to everyone. Google is first giving it to trusted cyber defenders through the Fairwind Program. Broader access will come later. Google says paid API customers and Google AI Ultra subscribers will be first in line. It has not announced a public release date.
Why the launch matters
The release follows a long gap at the top of Google’s model lineup. Google had been releasing smaller, cheaper Flash models. OpenAI and Anthropic were competing with their own frontier systems. Argon is Google’s attempt to compete at that highest level again.
The contest is also changing. A chatbot can answer one question well. A business assistant must handle documents, code, choices, and follow-up work. The value comes from keeping the task moving. That makes long workflows more important than a single clever answer.
Software work can require many files. Legal research can require many documents. Finance work can require careful comparisons. An AI must remember the goal across those steps. It must also stop when the information is uncertain.
What the benchmark table shows
Google’s published table reports 68.9% for Argon on the Vals Index knowledge-work test. GPT-6 Astra scored 63.1%. Claude Fable 5.1 scored 65.8%. Claude Opus 5.5 scored 67.0%.
On the DeepSWE v1.1 software-engineering test, Argon scored 77.9%. That was higher than the three comparison models. But Argon did not lead every test. GPT-6 Astra led FrontierSWE v2, and Claude Opus 5.5 led Terminal-Bench 4.0. These are controlled measurements, not guarantees about a real company’s results.
What can be confirmed
Google says Argon can find, validate, and patch critical software vulnerabilities. The claim is about defensive work. The company has not presented this as a reason to give the model unrestricted access.
Google also describes several safety measures. Argon is designed to refuse harmful requests. It is designed to resist indirect prompt injections. Google says it can monitor the model’s reasoning and actions, then stop execution when needed. Google is also hardening the isolated environments used for risky testing.
Reports on the launch, including Ars Technica’s account, also described a much larger output limit. Argon may produce up to one million tokens, compared with 64,000 in earlier Gemini models. More output can help with large tasks. It does not prove that every extra sentence is correct.
What remains unknown
The strongest evidence still comes from Google’s own benchmark table and early testing. Independent evaluators must check whether those results hold outside Google’s setup. Users also need to see whether Argon saves time, reduces mistakes, or lowers costs in real work.
The public release date is another unanswered question. Google says access will expand after more testing, but it has not set a schedule. The safety claims also need continued observation, especially when the model handles real tools and sensitive work.
What to watch next
The next signal will come from Fairwind testers and their cybersecurity results. After that, paid API users and Google AI Ultra subscribers should provide wider feedback. The key question is not whether Argon wins one table. It is whether the model can complete difficult work reliably, safely, and repeatedly.
Why Google’s new AI is built for long jobs
📰 Full story: Google unveils Gemini 4 Argon, but the new flagship is not open yet
Gemini 4 Argon is built to keep working through many steps.
benchmark
A fixed test that compares different AI systems.
Fairwind Program
Google’s safety testing program for a small group of users.
cybersecurity defense
Work that protects computers and networks.
💡 The gist
- Google introduced Gemini 4 Argon, a new AI model.
- It targets long jobs in code, business, and cyber defense.
- Most people cannot use it yet.
What makes it different?
Google DeepMind, Google’s AI research group, built Argon. The model targets jobs with many steps. Those jobs can include software, legal research, and finance. It also targets cybersecurity defense. That means protecting computers and networks.
A normal chatbot often answers one request. Argon is designed for longer workflows. It may plan steps, do work, and check results. This matters because business tasks rarely have one step. They often involve documents, code, and repeated decisions. Software work can require many files. Legal research can require many documents. Finance work can require careful comparisons. An AI must remember the goal across those steps. It must also stop when the information is uncertain.
What do the scores mean?
Google’s benchmark table gives Argon 68.9% on Vals Index. It gives Argon 77.9% on DeepSWE v1.1. Some rival models score lower there. Argon does not win every test. GPT-6 Astra and Claude Opus 5.5 lead some tests. A benchmark is a fixed test for comparing AI systems. A high score does not guarantee real workplace success.
Why is access limited?
Powerful AI can help people. It can also make larger mistakes. It may face harmful or unsafe uses. Google is testing Argon with trusted cyber defenders first. This testing uses the Fairwind Program. The program focuses on safe cybersecurity work. Google says paid API users and AI Ultra subscribers come later. It has not promised a public date.
Why does this matter?
If Argon works as promised, companies may save time. They may ask AI to handle longer assignments. People would still need to check important results. The model’s test scores are only one piece of evidence. Real work will provide the harder test.
What should we watch?
Independent testers should check Google’s claims. Developers should test Argon on real tasks. Safety reviews should continue before wider access. The key question is simple. Can Argon finish difficult work reliably and safely?
Gemini 4 Argon is a new computer helper
📰 Full story: Google unveils Gemini 4 Argon, but the new flagship is not open yet
Google showed an AI that can help with long jobs.
Google DeepMind
Google’s team that builds AI.
Gemini 4 Argon
Google’s new AI helper for difficult work.
cyber defense
Protecting computers from harm.
Google DeepMind (Google’s AI team) showed Gemini 4 Argon, a new computer helper.
Argon can help write computer instructions. It can help study law and money work. It can help protect computers. That job is called cyber defense.
Long jobs have many small steps. Argon tries to follow those steps. It can check what it did. That may help busy people.
Google says Argon did well on tests. A test is like a school check. A school check cannot show every skill.
Most people cannot use Argon yet. Safety testers will try it first. Google will watch its actions. Later, paying users may get access. Google has not chosen a day. People must see how it works in real jobs.