A systematization-of-knowledge paper by organizers and competitors analyzes AIxCC's design, the seven finalist architectures and results beyond the scoreboard. It reports that system stability and accuracy penalties decided rankings, that LLM-based systems found vulnerabilities a fuzzing baseline missed, and that among patches passing all automatic validation, manual review found semantic errors in 38-46% from baseline agents; the top two systems had 83.8% and 79.2% competition-scored patch accuracy.
Why it matters
It gives a detailed account, beyond the scoreboard, of what autonomous cyber reasoning systems achieved and where their patches failed.
Key facts
As stated in the sources, with where to find them.
- Team Atlanta scored 392.8 points vs 219.4 for second place; Theori's pre-penalty score exceeded Trail of Bits but a -16.3 accuracy penalty dropped it to third.Section 7.1
- A parallel-fuzzing baseline solved 34 of 63 challenge vulnerabilities (54%) but only 4 of 23 Java ones; six CRSs found 7-16 vulnerabilities the baseline missed.Section 7.3
- Competition-scored patch accuracy was 83.8% (Team Atlanta) and 79.2% (Trail of Bits); manual review of baseline-agent patches that passed all automatic checks found 37.7% (Claude Code) and 45.6% (MultiRetrieval) semantically incorrect.Table 10; Section 7.4 (KF 6-7)
- About 94% of LLM spending went to Anthropic and OpenAI models; no team exhausted its $50K LLM credit or $85K compute budget.Section 7.5, Table 8
- Companion site lists raw competition data (submission logs, traces, scoring breakdowns) as pending release.Companion site, artifacts list
- Accepted to USENIX Security 2026; v1 7 Feb 2026, v5 2 Aug 2026.arXiv listing
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
Sources
Related records
Aug 8, 2025
Mar 9, 2026
Aug 11, 2024
Sep 3, 2026
Aug 9, 2023
Mar 6, 2026