Chronicle/Defense & research

AIxCC SoK finds stability decided results and many validated AI patches were still semantically wrong

DefensePaperSignificance assistant-drafted

A systematization-of-knowledge paper by organizers and competitors analyzes AIxCC's design, the seven finalist architectures and results beyond the scoreboard. It reports that system stability and accuracy penalties decided rankings, that LLM-based systems found vulnerabilities a fuzzing baseline missed, and that among patches passing all automatic validation, manual review found semantic errors in 38-46% from baseline agents; the top two systems had 83.8% and 79.2% competition-scored patch accuracy.

Why it matters

It gives a detailed account, beyond the scoreboard, of what autonomous cyber reasoning systems achieved and where their patches failed.

Key facts

As stated in the sources, with where to find them.

  • Team Atlanta scored 392.8 points vs 219.4 for second place; Theori's pre-penalty score exceeded Trail of Bits but a -16.3 accuracy penalty dropped it to third.Section 7.1
  • A parallel-fuzzing baseline solved 34 of 63 challenge vulnerabilities (54%) but only 4 of 23 Java ones; six CRSs found 7-16 vulnerabilities the baseline missed.Section 7.3
  • Competition-scored patch accuracy was 83.8% (Team Atlanta) and 79.2% (Trail of Bits); manual review of baseline-agent patches that passed all automatic checks found 37.7% (Claude Code) and 45.6% (MultiRetrieval) semantically incorrect.Table 10; Section 7.4 (KF 6-7)
  • About 94% of LLM spending went to Anthropic and OpenAI models; no team exhausted its $50K LLM credit or $85K compute budget.Section 7.5, Table 8
  • Companion site lists raw competition data (submission logs, traces, scoring breakdowns) as pending release.Companion site, artifacts list
  • Accepted to USENIX Security 2026; v1 7 Feb 2026, v5 2 Aug 2026.arXiv listing

Findings that cite this record

Key questions this bears on

Sources

Related records