Findings/aixcc-systems-patched-most-found-bugs

In DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.

Qualifiedmeasured2 evidence records from 1 independent sourceassistant-drafted
Scope: what this does not show

Competition challenges; DARPA's own scoring.

Qualified: Still standing, but later work narrows how far it applies.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Status history

  1. 2024-08-11ReportedSemifinal results. · record
  2. 2026-02-07QualifiedA later review finds 16-21% of top patches semantically wrong. · record
  3. 2026-09-25QualifiedcorrectionThe 16-21% figure is competition-scored submission accuracy and does not reduce DARPA's 43 counted patches. The qualification now rests on PatchBench: agents from top AIxCC teams lose much of their solve rate under stronger-than-crash validation. · record