Google DeepMind announced Gemini 4 Argon on 2026-09-30 and says it is rolling out to a set of trusted cyber defenders through the Fairwind Program, and says it will release the model without cyber guardrails for those defenders and Google's internal teams. Google claims the model can autonomously find, validate and patch critical vulnerabilities, ties for first on CWE-bench v1 at 68%, and improves on Gemini 3.8 Flash Cyber in vulnerability discovery, without giving numbers for the comparison. It says broad release awaits further safeguard work in four areas (misuse, prompt injection, misalignment and hardening), citing its Frontier Safety Framework for the misuse safeguards; all of these are Google-reported claims that have not been independently verified.
It announces plans to extend the gated, guardrail-reduced access model Google used for Gemini 3.8 Flash Cyber to its next frontier model, with the access terms and safeguard evidence left largely to the vendor's own description.
Key facts
As stated in the sources, with where to find them.
- Access: Argon is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google says it will release the model without cyber guardrails for trusted defenders and its own internal teams. The Argon post states no eligibility terms; the Gemini 3.8 Flash Cyber post describes Fairwind as providing trusted government authorities, critical infrastructure operators and software maintainers with prioritized access to Gemini 3.8 Flash Cyber and an application route; whether the same terms apply to Argon is not stated.Argon post, intro and Leading in defensive cybersecurity; 3.8 post, Gemini 3.8 Flash and Cyber: get started
- Phasing: Google describes a phased approach, says it is engaged in the US government's voluntary process for pre-release model access, and says it will iterate on guardrails with early testers before making Argon available to developers, enterprises and consumers, starting with paid API customers and Google AI Ultra subscribers.Argon post, intro; closing section on rollout
- Google says Argon was trained to be highly capable at cybersecurity defense and can autonomously find, validate and patch critical software vulnerabilities. It states Wiz is using Argon through its Scan for Good initiative for free protection of critical public infrastructure, and that in an early demonstration the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide that previous frontier models had missed. The post gives no vulnerability identifier, affected vendor or disclosure status.Argon post, Leading in defensive cybersecurity
- CWE-bench v1 (vulnerability remediation): Google says Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber's performance on CWE-bench v0. The post does not name the other top-scoring model, give a 3.8 Flash Cyber score on v1, or say who ran the benchmark. The 3.8 Flash Cyber post reported 47.2% pass@1 on a CWE-Bench run by Collinear without naming a version; because the Argon post refers to 3.8 Flash Cyber's performance on v0, the two scores are probably on different versions, which is this record's inference.Argon post, Leading in defensive cybersecurity; 3.8 post, Automated patching
- Comparison with 3.8 Flash Cyber: Google reports that on its internal vulnerability benchmark across 20 programming languages Argon uncovered a wide range of exposures, and that on Wiz's internal black-box penetration-testing benchmark Argon outperforms 3.8 Flash Cyber in discovering attack surface, identifying vulnerabilities and producing proof-of-concept evidence. No scores are given; the 3.8 post reported a success rate exceeding 70% on its internal 20-language benchmark, which the Argon post does not restate.Argon post, Leading in defensive cybersecurity; 3.8 post, Autonomous vulnerability discovery
- Misuse safeguards: Google says the model is designed to refuse harmful cyber and CBRN requests while preserving legitimate dual-use research, as per its Frontier Safety Framework; that it is strengthening safeguard robustness for this launch, including better monitoring of the model's internal activations; and that the safeguards were tested by internal and external red teams using manual and automated attacks. The post gives no results, red-team counts or capability-threshold determination.Argon post, Strengthening frontier safeguards, Defending against misuse
- Other safeguards: Google calls Argon its most resilient model against indirect prompt injection and says it leads on Gray Swan's indirect prompt injection benchmark (no score given); says it monitors the model's chain of thought and actions and stops execution when necessary, using a similar system on training runs with alerts to an incident response team and without feeding findings back into training; and says it is isolating and sealing sandboxed environments before high-risk training or evaluations.Argon post, Strengthening frontier safeguards
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.