Methods/Attack technique

AI-assisted vulnerability exploitation

Using AI models or agents to find vulnerabilities and turn them into working exploits.

14 records1 attack3 capability3 defense7 policy3 findings (2 measured)First recorded 2023-10assistant-drafted

How it works

Models reason over code, crashes, and documentation to locate flaws and assemble exploitation steps. Benchmarks measure how far along that path they get; threat reports describe criminal use.

What we know

2 reported, 1 qualified

Records over time

RangeLanes
12 of 14 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
202420252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

12 records · newest first
Sep 2026
Sep 2, 2026
Google releases Gemini 3.8 Flash Cyber for trusted defenders, emphasizing automated patching
DefenseTool releaseGoogle, Google DeepMind, Collinear

Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-tuned model with more permissive cyber mitigations, available only to trusted defenders through a new Fairwind Program. Google says it prioritized vulnerability fixing over exploitation and reports 47.2% pass@1 on Collinear's CWE-Bench patching benchmark, over 70% on an internal 20-language discovery benchmark, and 2.6 times more correct Chrome patches than larger commercial models.

Jul 2026
Jul 23, 2026
UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models
CapabilityEvaluation reportUK AI Security Institute, US Center for AI Standards and Innovation, Moonshot AI

The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.

Jul 2, 2026
Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners
PolicyFrameworkAnthropic

Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response.

Jun 2026
Jun 30, 2026
US lifts export controls on Fable 5 and Mythos 5; Anthropic redeploys with new cyber classifier
PolicyRegulationAnthropic, US Department of Commerce, US Center for AI Standards and Innovation

Anthropic announced that export controls on Fable 5 and Mythos 5 had been lifted and that Fable 5 would be redeployed globally from July 1, 2026 with an improved safety classifier. Anthropic says the classifier blocks the technique described in an Amazon report in over 99% of cases and that CAISI researchers tested its prior and new safeguards. Mythos 5 access was restored for a set of US organizations after government approval on June 26.

May 2026
May 18, 2026
Maintainers report AI-generated vulnerability reports overwhelming kernel and bounty triage
DefenseIncidentLinux kernel maintainers, curl project, GitHub

Help Net Security reported that Linus Torvalds described the Linux kernel security list as almost entirely unmanageable because of heavily duplicated AI-assisted reports, and that GitHub tightened its bug bounty submission requirements, with a GitHub engineer saying some programs elsewhere had shut down. The article also notes that curl ended bounty payments after a surge of low-quality AI reports.

May 13, 2026
ExploitBench grades AI exploit development as a 16-step capability ladder on V8 bugs
CapabilityBenchmarkCarnegie Mellon University, Bugcrowd

Carnegie Mellon researchers released ExploitBench, which scores exploitation progress on 41 V8 JavaScript-engine vulnerabilities across 16 flags from reaching the bug through arbitrary read/write, control-flow hijack and code execution. The paper reports that public models routinely reach and crash vulnerable code but rarely achieve arbitrary code execution, while one private frontier model succeeded on roughly half of cases.

May 11, 2026
ExploitGym benchmark measures whether AI agents can turn real vulnerabilities into working exploits
CapabilityBenchmarkUC Berkeley, Anthropic, OpenAI

Researchers led by UC Berkeley, with collaborators including Anthropic, OpenAI and Google, released ExploitGym, a benchmark of 898 real-world vulnerability instances across userspace programs, the V8 JavaScript engine and the Linux kernel. Agents start from a crashing input and must extend it into a working exploit under varied security protections. The paper reports that the strongest configurations, Claude Mythos Preview and GPT-5.5, produced working exploits for 157 and 120 instances respectively.

May 11, 2026
Google Threat Intelligence reports the first criminal zero-day exploit it believes was AI-developed, disrupted before planned mass use
AttackMisuse reportGoogle Threat Intelligence Group, UNC6780 (TeamPCP)

Google Threat Intelligence Group reported that cybercriminals planned a mass-exploitation campaign using a two-factor-authentication bypass in an open-source web administration tool, and assessed with high confidence that an AI model supported discovery and weaponization of the flaw. GTIG worked with the vendor on disclosure and disrupted the activity. The same report describes PRC-nexus actors using agentic frameworks such as Hexstrike and Strix for reconnaissance and vulnerability validation, and Android malware (PROMPTSPY) that calls Gemini to drive the device UI.

Jul 2025
Jul 15, 2025
Google says Big Sleep found SQLite CVE-2025-6965 before attackers could exploit it
DefenseVulnerability disclosureGoogle, Google DeepMind, Google Project Zero

Google reports that, working from Google Threat Intelligence information, the Big Sleep agent found a critical SQLite memory-corruption flaw (CVE-2025-6965) that Google says was known only to threat actors and at risk of exploitation. Google says it reported the flaw for patching before attackers could exploit it, says it believes this is the first time an AI agent directly foiled an in-the-wild exploitation effort, and says Big Sleep is being applied to open-source projects.

Jul 10, 2025
EU GPAI Code of Practice Safety and Security chapter lists cyber offence as a specified systemic risk
PolicyFrameworkEuropean Commission

The European Commission received the final General-Purpose AI Code of Practice, whose Safety and Security chapter applies to providers of models with systemic risk under Article 55 of the AI Act. The chapter treats cyber offence as one of four specified systemic risks, requires a security goal covering non-state external and insider threats, and sets serious incident reporting deadlines that include five days for serious cybersecurity breaches.

May 2025
May 7, 2025
UK NCSC judges AI-assisted vulnerability research is the most significant AI cyber development to 2027
PolicyGuidanceUK National Cyber Security Centre

The NCSC's second assessment judges that AI will almost certainly make elements of intrusion more effective through 2027, with AI-assisted vulnerability research and exploit development the most significant development. It warns that the window between disclosure and exploitation, already days, will shrink further, and judges fully automated end-to-end advanced attacks unlikely before 2027.

Apr 2025
Apr 15, 2025
OpenAI Preparedness Framework v2 sets High and Critical cybersecurity capability thresholds
PolicyFrameworkOpenAI

OpenAI's Preparedness Framework version 2 makes cybersecurity one of three Tracked Categories and defines High and Critical capability thresholds, each tied to required safeguards. High covers automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities; Critical covers autonomous zero-day development across many hardened critical systems, and at Critical OpenAI commits to halt further development until adequate safeguards are specified.

All records