<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Vulnerability repair · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/vulnerability-repair/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/vulnerability-repair/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on vulnerability repair, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>Correction to a finding (reconfirmed as corroborated): Checking only that the original crash no longer reproduces overstates how often AI-generated patches actually fix the vulnerability.</title>
<link>https://agentic-cyber-explorer.pages.dev/findings/crash-checks-overstate-patch-success/</link>
<guid isPermaLink="false">correction:crash-checks-overstate-patch-success:2026-09-25:corroborated</guid>
<pubDate>Fri, 25 Sep 2026 12:00:00 GMT</pubDate>
<category>Correction</category>
<description>Correction: the review paper's support is its manual review of baseline agents (38-46% of fully validated patches semantically wrong), not the competition-scored accuracy figures.</description>
</item>
<item>
<title>Correction to a finding (reconfirmed as qualified): In DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.</title>
<link>https://agentic-cyber-explorer.pages.dev/findings/aixcc-systems-patched-most-found-bugs/</link>
<guid isPermaLink="false">correction:aixcc-systems-patched-most-found-bugs:2026-09-25:qualified</guid>
<pubDate>Fri, 25 Sep 2026 12:00:00 GMT</pubDate>
<category>Correction</category>
<description>Correction: the 16-21% figure is competition-scored submission accuracy and does not reduce DARPA's 43 counted patches. The qualification now rests on PatchBench: agents from top AIxCC teams lose much of their solve rate under stronger-than-crash validation.</description>
</item>
<item>
<title>PatchBench finds PoC-only checks inflate AI patching success 1.83x and 25% of patches look memorized</title>
<link>https://agentic-cyber-explorer.pages.dev/events/patchbench-vulnerability-patching-validity-2026/</link>
<guid isPermaLink="false">event:patchbench-vulnerability-patching-validity-2026</guid>
<pubDate>Thu, 03 Sep 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>PatchBench, from the University of Maryland's AI Security Lab, evaluates 11 patching agents, including the top three AIxCC systems, on 213 C/C++ tasks whose true fixes lie outside the crash stack, using vulnerability transplant and code mutation to limit memorization. It finds that accepting a patch because the original proof-of-concept no longer crashes inflates solve rates by 1.83x on average, and that about 25% of agent patches closely resemble historical developer fixes. It directly challenges how AI vulnerability-repair results, including competition results, are validated.</description>
</item>
<item>
<title>Google releases Gemini 3.8 Flash Cyber for trusted defenders, emphasizing automated patching</title>
<link>https://agentic-cyber-explorer.pages.dev/events/google-gemini-3-8-flash-cyber-2026/</link>
<guid isPermaLink="false">event:google-gemini-3-8-flash-cyber-2026</guid>
<pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-tuned model with more permissive cyber mitigations, available only to trusted defenders through a new Fairwind Program. Google says it prioritized vulnerability fixing over exploitation and reports 47.2% pass@1 on Collinear's CWE-Bench patching benchmark, over 70% on an internal 20-language discovery benchmark, and 2.6 times more correct Chrome patches than larger commercial models. It is a gated, defense-oriented model release that foregrounds patching metrics rather than offensive capability.</description>
</item>
<item>
<title>European Commission presents EU Action Plan on Cybersecurity and Artificial Intelligence</title>
<link>https://agentic-cyber-explorer.pages.dev/events/eu-action-plan-cybersecurity-ai-2026/</link>
<guid isPermaLink="false">event:eu-action-plan-cybersecurity-ai-2026</guid>
<pubDate>Tue, 07 Jul 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The Commission presented an action plan responding to advanced AI models that can both improve and undermine cybersecurity. It plans an EU capacity to evaluate AI models, a European blueprint for structured access to advanced AI capabilities developed with ENISA, a secure ENISA-JRC platform to test AI for cybersecurity, AI-assisted vulnerability fixing, and a campaign to secure critical open-source software. ENISA published its own recommendations for the frontier AI era the same day. It is a dedicated EU policy response to frontier AI cyber capability, including structured access for defenders.</description>
</item>
<item>
<title>Five Eyes cyber agency heads tell leaders AI is shifting cyber risk on a timescale of months</title>
<link>https://agentic-cyber-explorer.pages.dev/events/five-eyes-ai-shift-in-cyber-risk-statement-2026/</link>
<guid isPermaLink="false">event:five-eyes-ai-shift-in-cyber-risk-statement-2026</guid>
<pubDate>Mon, 22 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The heads of the Five Eyes cyber agencies issued a joint statement that AI is rapidly transforming cyber risk and that organizations must act within months, not years. They ask leaders to reduce attack surface, accelerate patching as exploitation windows shorten, replace unsupported legacy systems, strengthen identity controls, and prepare for incidents. It is the highest-level joint government signal that frontier AI vulnerability discovery changes patching expectations.</description>
</item>
<item>
<title>Executive Order 14409 creates classified cyber benchmarking for covered frontier models and a clearinghouse</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-eo-14409-frontier-ai-cyber-benchmarking-2026/</link>
<guid isPermaLink="false">event:us-eo-14409-frontier-ai-cyber-benchmarking-2026</guid>
<pubDate>Tue, 02 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Executive Order 14409 directs Treasury, NSA and CISA to develop a classified benchmarking process to assess advanced cyber capabilities of AI models and designate covered frontier models, with a voluntary framework for pre-release government and trusted-partner access. It also orders an AI cybersecurity clearinghouse to coordinate vulnerability scanning, validation and remediation with industry, and states it does not create mandatory licensing or pre-clearance. It is a US mechanism that designates models by cyber capability and gives the government early access before release to other trusted partners.</description>
</item>
<item>
<title>Glasswing update: over 10,000 high-severity bugs found, but only 75 of 530 disclosed OSS bugs patched</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-glasswing-initial-update-2026/</link>
<guid isPermaLink="false">event:anthropic-glasswing-initial-update-2026</guid>
<pubDate>Fri, 22 May 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic reports that about 50 Glasswing partners used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities, and that its own scan of over 1,000 open-source projects produced 6,202 model-estimated high/critical findings. Of 1,752 assessed, mostly by six independent firms, 90.6% were true positives; Anthropic estimates 530 high/critical bugs disclosed, of which 75 were patched, and says triage and patching capacity, not discovery, is the bottleneck. It gives rare pipeline-level numbers showing AI vulnerability discovery outpacing the human capacity to verify, disclose and fix.</description>
</item>
<item>
<title>OSS-CRS makes AIxCC reasoning systems runnable locally; OpenSSF adopts it as a sandbox project</title>
<link>https://agentic-cyber-explorer.pages.dev/events/oss-crs-aixcc-systems-openssf-2026/</link>
<guid isPermaLink="false">event:oss-crs-aixcc-systems-openssf-2026</guid>
<pubDate>Mon, 09 Mar 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Researchers led by Georgia Tech released OSS-CRS, a locally deployable framework for running and combining AIxCC cyber reasoning systems, noting that all seven open-sourced finalist systems depended on competition cloud infrastructure that no longer exists. Porting the winning Atlantis system, they found 10 previously unknown bugs (three high severity) in 8 OSS-Fuzz projects; OpenSSF welcomed OSS-CRS into its AI/ML Security Working Group in April 2026. It addresses the gap between open-sourcing competition systems and making them usable by maintainers.</description>
</item>
<item>
<title>OpenAI relaunches Aardvark as Codex Security, reporting 1.2M commits scanned and 14 CVEs</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-codex-security-research-preview-2026/</link>
<guid isPermaLink="false">event:openai-codex-security-research-preview-2026</guid>
<pubDate>Fri, 06 Mar 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>OpenAI renamed Aardvark to Codex Security and opened a research preview to ChatGPT Pro, Enterprise, Business and Edu customers. OpenAI reports that in 30 days it scanned over 1.2 million commits in its beta cohort and flagged 792 critical and 10,561 high-severity findings, that beta changes cut false positives by more than 50%, and that its open-source reports led to 14 CVEs. It gives rare operational-scale figures, self-reported by the vendor, on AI code-scanning volume and false-positive reduction, alongside a program for open-source maintainers.</description>
</item>
<item>
<title>Anthropic releases Claude Code Security in limited preview to scan code and propose patches</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-claude-code-security-2026/</link>
<guid isPermaLink="false">event:anthropic-claude-code-security-2026</guid>
<pubDate>Fri, 20 Feb 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic released Claude Code Security as a limited research preview for Enterprise and Team customers, with expedited free access for open-source maintainers. The tool reasons about data flow across a codebase, re-examines each finding in a multi-stage verification pass, assigns severity and confidence ratings, and proposes patches that are applied only with human approval. It packages frontier-model vulnerability finding for defenders with explicit human approval gates, amid concerns about the same capability aiding attackers.</description>
</item>
<item>
<title>AIxCC SoK finds stability decided results and many validated AI patches were still semantically wrong</title>
<link>https://agentic-cyber-explorer.pages.dev/events/aixcc-sok-competition-lessons-2026/</link>
<guid isPermaLink="false">event:aixcc-sok-competition-lessons-2026</guid>
<pubDate>Sat, 07 Feb 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>A systematization-of-knowledge paper by organizers and competitors analyzes AIxCC's design, the seven finalist architectures and results beyond the scoreboard. It reports that system stability and accuracy penalties decided rankings, that LLM-based systems found vulnerabilities a fuzzing baseline missed, and that among patches passing all automatic validation, manual review found semantic errors in 38-46% from baseline agents; the top two systems had 83.8% and 79.2% competition-scored patch accuracy. It gives a detailed account, beyond the scoreboard, of what autonomous cyber reasoning systems achieved and where their patches failed.</description>
</item>
<item>
<title>Anthropic reports over 500 human-validated high-severity open-source vulnerabilities found with Claude Opus 4.6</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-opus-4-6-500-zero-days-2026/</link>
<guid isPermaLink="false">event:anthropic-opus-4-6-500-zero-days-2026</guid>
<pubDate>Thu, 05 Feb 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic's Frontier Red Team reports that Claude Opus 4.6, run in a VM with standard tools but no custom harness, found and validated more than 500 high-severity vulnerabilities in open-source software, focusing on memory corruption that can be confirmed with sanitizers. Every bug was validated before reporting, initially by Anthropic researchers who also wrote patches and later with external researchers; examples include Ghostscript, OpenSC and CGIF. It shows a general-purpose model finding bugs in heavily fuzzed code out of the box and describes the validation effort needed to avoid burdening maintainers.</description>
</item>
<item>
<title>OpenAI announces Aardvark, a GPT-5 agent that finds, validates and proposes patches for vulnerabilities</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-aardvark-private-beta-2025/</link>
<guid isPermaLink="false">event:openai-aardvark-private-beta-2025</guid>
<pubDate>Thu, 30 Oct 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>OpenAI announced Aardvark, a GPT-5-powered agent in private beta that builds a threat model of a repository, scans commits, tries to trigger suspected flaws in a sandbox, and attaches Codex-generated patches for human review. OpenAI reports 92% recall on known and synthetically introduced vulnerabilities in its 'golden' repositories and ten CVEs from open-source scanning, and planned pro-bono scanning for some non-commercial projects. It combined LLM reasoning, sandbox validation and patch generation in one defensive agent from a frontier lab, later relaunched as Codex Security.</description>
</item>
<item>
<title>Google DeepMind introduces CodeMender, an agent that patches and hardens code, with 72 upstreamed fixes</title>
<link>https://agentic-cyber-explorer.pages.dev/events/google-deepmind-codemender-2025/</link>
<guid isPermaLink="false">event:google-deepmind-codemender-2025</guid>
<pubDate>Mon, 06 Oct 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Google DeepMind introduced CodeMender, an agent built on Gemini Deep Think models that combines static and dynamic analysis, fuzzing, differential testing and SMT solvers with LLM-based critique to generate and validate security patches. DeepMind reports 72 security fixes upstreamed to open-source projects over six months, all reviewed by human researchers before submission; in May 2026 Google said it would fold CodeMender into its enterprise agent platform. It is a leading example of an AI agent aimed at the repair side of vulnerability management, including proactive rewriting to remove bug classes.</description>
</item>
<item>
<title>Anthropic says it trained Claude Sonnet 4.5 for defensive vulnerability finding and patching</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-building-ai-cyber-defenders-2025/</link>
<guid isPermaLink="false">event:anthropic-building-ai-cyber-defenders-2025</guid>
<pubDate>Fri, 03 Oct 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic reports that a small team focused Claude Sonnet 4.5 training on finding and patching vulnerabilities and on testing simulated security infrastructure, while avoiding enhancements that clearly favour offence. It reports Sonnet 4.5 results on Cybench and CyberGym, a preliminary patching study in which 15% of patches were judged semantically equivalent to human references, and invites work on SOC and SIEM automation. It is an explicit statement by a frontier lab that it steered model training toward defensive cyber skills, with measured results and patching caveats.</description>
</item>
<item>
<title>AIxCC final: Team Atlanta wins as systems patch 43 of 54 found synthetic bugs and find 18 real ones</title>
<link>https://agentic-cyber-explorer.pages.dev/events/darpa-aixcc-final-results-2025/</link>
<guid isPermaLink="false">event:darpa-aixcc-final-results-2025</guid>
<pubDate>Fri, 08 Aug 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>DARPA reports that seven finalist cyber reasoning systems analyzed over 54 million lines of code, found 54 unique synthetic vulnerabilities in 63 challenges and patched 43, and found 18 real non-synthetic vulnerabilities with 11 patches. Team Atlanta won $4 million, Trail of Bits $3 million and Theori $1.5 million; DARPA and ARPA-H added $1.4 million for real-world integration and four systems were open-sourced on the day. It is an organizer-verified, competition-scale measurement of autonomous AI vulnerability discovery and patching, with open-sourced systems others can reuse.</description>
</item>
<item>
<title>SEC-bench automatically builds real vulnerability tasks and finds agents patch at most 34%</title>
<link>https://agentic-cyber-explorer.pages.dev/events/sec-bench-2025/</link>
<guid isPermaLink="false">event:sec-bench-2025</guid>
<pubDate>Fri, 13 Jun 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>SEC-bench uses multi-agent scaffolding to construct reproducible vulnerability instances with test environments and validated patches from real projects, at about $0.87 per instance. The authors report that LLM agents reached at most 18.0% on proof-of-concept generation and 34.0% on vulnerability patching. It offers a cheaper route to fresh vulnerability benchmarks and shows low agent patching rates even with call-stack hints and a build-and-PoC check.</description>
</item>
<item>
<title>BountyBench measures AI agents on detect, exploit and patch tasks from real bug bounties</title>
<link>https://agentic-cyber-explorer.pages.dev/events/stanford-bountybench-2025/</link>
<guid isPermaLink="false">event:stanford-bountybench-2025</guid>
<pubDate>Wed, 21 May 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>BountyBench, from Stanford-led researchers, builds 40 bug bounties across 25 real-world systems into 120 Detect, Exploit and Patch tasks with dollar values attached. In the first version the best Detect score was 5%, while OpenAI Codex CLI and Claude Code scored 90% and 87.5% on Patch, well above their Exploit scores. A July 2025 revision with more agents reported Codex CLI with o3-high at 12.5% on Detect and 90% on Patch. It puts offensive and defensive agent performance on the same real codebases and expresses results in bounty dollars.</description>
</item>
<item>
<title>Meta releases AutoPatchBench to test AI repair of fuzzing-found C/C++ vulnerabilities</title>
<link>https://agentic-cyber-explorer.pages.dev/events/meta-autopatchbench-2025/</link>
<guid isPermaLink="false">event:meta-autopatchbench-2025</guid>
<pubDate>Tue, 29 Apr 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Meta introduced AutoPatchBench, part of CyberSecEval 4, with 136 fuzzing-identified C/C++ vulnerabilities and verified fixes, plus a 113-case Lite subset with single-function root causes. Patches are checked by build and crash reproduction, then fuzzing and white-box differential testing; Meta's reference agent generated crash-stopping patches in about 60% of cases, but only 5-11% passed the stricter checks. It showed early that crash-only acceptance greatly overstates how often AI-generated security patches are actually correct.</description>
</item>
<item>
<title>AIxCC semifinal: AI systems find 22 synthetic vulnerabilities, patch 15, and find one real SQLite bug</title>
<link>https://agentic-cyber-explorer.pages.dev/events/darpa-aixcc-semifinal-results-2024/</link>
<guid isPermaLink="false">event:darpa-aixcc-semifinal-results-2024</guid>
<pubDate>Sun, 11 Aug 2024 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>DARPA reports that in the AIxCC semifinal at DEF CON 32, nearly 40 cyber reasoning systems were tested on challenge projects based on Jenkins, the Linux kernel, Nginx, SQLite3 and Apache Tika. Competitors' systems found 22 unique synthetic vulnerabilities, patched 15, and found one real-world SQLite3 bug; seven teams advanced with $2 million each and must open-source their systems after the final. It gave organizer-verified numbers on how well AI cyber reasoning systems could find and patch vulnerabilities in challenge projects built on widely used open-source software.</description>
</item>
<item>
<title>ARVO dataset makes OSS-Fuzz vulnerabilities reproducible with located fixes (over 5,000 at release, 6,100+ by 2026)</title>
<link>https://agentic-cyber-explorer.pages.dev/events/arvo-reproducible-vulnerability-dataset-2024/</link>
<guid isPermaLink="false">event:arvo-reproducible-vulnerability-dataset-2024</guid>
<pubDate>Sun, 04 Aug 2024 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>ARVO (Atlas of Reproducible Vulnerabilities for Open Source Software) builds reproducible vulnerability cases from OSS-Fuzz, each with a triggering input, a rebuildable environment and an automatically located fixing patch. The August 2024 first version reported over 5,000 memory vulnerabilities across 250+ C/C++ projects; the authors' June 2026 revision reports over 6,100 vulnerabilities across 311 projects, 81% reproduction success and 89.4% accuracy on located patches. The paper is accepted at IEEE EuroS&amp;P 2026. Reproducible vulnerability/fix pairs are the raw material for evaluating AI repair agents, and ARVO underlies several later benchmarks.</description>
</item>
<item>
<title>Google reports Gemini-based pipeline fixed 15% of sanitizer bugs found in its unit tests</title>
<link>https://agentic-cyber-explorer.pages.dev/events/google-ai-powered-patching-2024/</link>
<guid isPermaLink="false">event:google-ai-powered-patching-2024</guid>
<pubDate>Wed, 31 Jan 2024 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>A Google Security Engineering technical report describes an automated pipeline that reproduces sanitizer-detected bugs in C/C++, Java and Go, prompts an LLM for fixes, tests them, and surfaces the best candidate for human review. Google reports that Gemini fixed 15% of sanitizer bugs discovered during unit tests, resulting in hundreds of patches. It is an early production-scale data point on LLM-generated security fixes with human review, a model later extended by CodeMender.</description>
</item>
<item>
<title>DARPA launches the AI Cyber Challenge to build AI systems that find and fix open-source vulnerabilities</title>
<link>https://agentic-cyber-explorer.pages.dev/events/darpa-aixcc-launch-2023/</link>
<guid isPermaLink="false">event:darpa-aixcc-launch-2023</guid>
<pubDate>Wed, 09 Aug 2023 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>At Black Hat USA 2023, DARPA announced the AI Cyber Challenge (AIxCC), a two-year competition to build AI-driven systems that automatically find and fix vulnerabilities in critical open-source software. Anthropic, Google, Microsoft and OpenAI agreed to provide technology and expertise to competitors, OpenSSF served as challenge advisor, and semifinal and final rounds were scheduled for DEF CON 2024 and 2025. AIxCC became a large public test of LLM-based cyber reasoning systems for defensive vulnerability discovery and repair.</description>
</item>
</channel>
</rss>
