Findings/crash-checks-overstate-patch-success

Checking only that the original crash no longer reproduces overstates how often AI-generated patches actually fix the vulnerability.

Corroboratedmeasured2 evidence records from 2 independent sourcesassistant-drafted
Scope: what this does not show

Two analyses of research and competition patches. PatchBench's 1.83x is measured on tasks selected so that the true fix lies outside the crash stack, so it is not a base rate.

Corroborated: Supported by at least two independent sources.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2026-02-07ReportedAIxCC review finds semantically wrong top patches. · record
  2. 2026-09-03CorroboratedPatchBench measures the inflation directly. · record
  3. 2026-09-25CorroboratedcorrectionThe review paper's support is its manual review of baseline agents (38-46% of fully validated patches semantically wrong), not the competition-scored accuracy figures. · record