Scope: what this does not show
Two analyses of research and competition patches. PatchBench's 1.83x is measured on tasks selected so that the true fix lies outside the crash stack, so it is not a base rate.
Corroborated: Supported by at least two independent sources.
Evidence
Feb 7, 2026
AIxCC SoK finds stability decided results and many validated AI patches were still semantically wrong
Manual review found 37.7% and 45.6% of baseline-agent patches that passed all automatic checks semantically incorrect.
Sep 3, 2026
PatchBench finds PoC-only checks inflate AI patching success 1.83x and 25% of patches look memorized
Proof-of-concept-only checks inflated success 1.83 times on PatchBench's tasks.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies In DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.The AIxCC teams' agents lose much of their solve rate under stronger validation on PatchBench tasks; the competition's counted patches were not re-validated.
Key questions that rely on this finding
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
Status history
- 2026-02-07ReportedAIxCC review finds semantically wrong top patches. · record
- 2026-09-03CorroboratedPatchBench measures the inflation directly. · record
- 2026-09-25CorroboratedcorrectionThe review paper's support is its manual review of baseline agents (38-46% of fully validated patches semantically wrong), not the competition-scored accuracy figures. · record