<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Exploit development · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/exploit-development/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/exploit-development/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on exploit development, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>Correction to a finding (reconfirmed as qualified): On ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.</title>
<link>https://agentic-cyber-explorer.pages.dev/findings/frontier-models-produce-working-exploits/</link>
<guid isPermaLink="false">correction:frontier-models-produce-working-exploits:2026-09-25:qualified</guid>
<pubDate>Fri, 25 Sep 2026 12:00:00 GMT</pubDate>
<category>Correction</category>
<description>Correction: the pipeline audit covered eight knowledge and multiple-choice benchmarks, not ExploitGym. The qualification rests on ExploitBench, where no publicly deployed model reached code execution on V8.</description>
</item>
<item>
<title>UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models</title>
<link>https://agentic-cyber-explorer.pages.dev/events/aisi-caisi-kimi-k3-cyber-assessment-2026/</link>
<guid isPermaLink="false">event:aisi-caisi-kimi-k3-cyber-assessment-2026</guid>
<pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development. It is an example of the two governments jointly evaluating a foreign open-weight model's cyber capability within a week of release.</description>
</item>
<item>
<title>ExploitBench grades AI exploit development as a 16-step capability ladder on V8 bugs</title>
<link>https://agentic-cyber-explorer.pages.dev/events/exploitbench-benchmark-2026/</link>
<guid isPermaLink="false">event:exploitbench-benchmark-2026</guid>
<pubDate>Wed, 13 May 2026 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>Carnegie Mellon researchers released ExploitBench, which scores exploitation progress on 41 V8 JavaScript-engine vulnerabilities across 16 flags from reaching the bug through arbitrary read/write, control-flow hijack and code execution. The paper reports that public models routinely reach and crash vulnerable code but rarely achieve arbitrary code execution, while one private frontier model succeeded on roughly half of cases. Graded scoring separates reaching or crashing a bug from building a working exploit, which crash-as-success benchmarks conflate.</description>
</item>
<item>
<title>Google Threat Intelligence reports the first criminal zero-day exploit it believes was AI-developed, disrupted before planned mass use</title>
<link>https://agentic-cyber-explorer.pages.dev/events/gtig-ai-developed-zero-day-2026/</link>
<guid isPermaLink="false">event:gtig-ai-developed-zero-day-2026</guid>
<pubDate>Mon, 11 May 2026 12:00:00 GMT</pubDate>
<category>Attacks &amp; incidents</category>
<description>Google Threat Intelligence Group reported that cybercriminals planned a mass-exploitation campaign using a two-factor-authentication bypass in an open-source web administration tool, and assessed with high confidence that an AI model supported discovery and weaponization of the flaw. GTIG worked with the vendor on disclosure and disrupted the activity. The same report describes PRC-nexus actors using agentic frameworks such as Hexstrike and Strix for reconnaissance and vulnerability validation, and Android malware (PROMPTSPY) that calls Gemini to drive the device UI. GTIG calls it the first identified instance of a zero-day exploit it believes was AI-developed by cybercrime actors.</description>
</item>
<item>
<title>ExploitGym benchmark measures whether AI agents can turn real vulnerabilities into working exploits</title>
<link>https://agentic-cyber-explorer.pages.dev/events/exploitgym-benchmark-2026/</link>
<guid isPermaLink="false">event:exploitgym-benchmark-2026</guid>
<pubDate>Mon, 11 May 2026 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>Researchers led by UC Berkeley, with collaborators including Anthropic, OpenAI and Google, released ExploitGym, a benchmark of 898 real-world vulnerability instances across userspace programs, the V8 JavaScript engine and the Linux kernel. Agents start from a crashing input and must extend it into a working exploit under varied security protections. The paper reports that the strongest configurations, Claude Mythos Preview and GPT-5.5, produced working exploits for 157 and 120 instances respectively. ExploitGym became a shared exploit-development yardstick in 2026 lab system cards and was the evaluation running during the Hugging Face intrusion.</description>
</item>
<item>
<title>UK NCSC judges AI-assisted vulnerability research is the most significant AI cyber development to 2027</title>
<link>https://agentic-cyber-explorer.pages.dev/events/ncsc-ai-cyber-threat-to-2027-2025/</link>
<guid isPermaLink="false">event:ncsc-ai-cyber-threat-to-2027-2025</guid>
<pubDate>Wed, 07 May 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The NCSC's second assessment judges that AI will almost certainly make elements of intrusion more effective through 2027, with AI-assisted vulnerability research and exploit development the most significant development. It warns that the window between disclosure and exploitation, already days, will shrink further, and judges fully automated end-to-end advanced attacks unlikely before 2027. It is a government forecast on autonomous attack timelines that 2026 frontier model evidence can be tested against.</description>
</item>
<item>
<title>OpenAI Preparedness Framework v2 sets High and Critical cybersecurity capability thresholds</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-preparedness-framework-v2-2025/</link>
<guid isPermaLink="false">event:openai-preparedness-framework-v2-2025</guid>
<pubDate>Tue, 15 Apr 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>OpenAI's Preparedness Framework version 2 makes cybersecurity one of three Tracked Categories and defines High and Critical capability thresholds, each tied to required safeguards. High covers automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities; Critical covers autonomous zero-day development across many hardened critical systems, and at Critical OpenAI commits to halt further development until adequate safeguards are specified. Its cyber thresholds are explicitly about autonomous, tool-augmented operation, so they are the operative gate for OpenAI's agentic cyber models in 2025-2026.</description>
</item>
</channel>
</rss>
