A systematization-of-knowledge paper by organizers and competitors analyzes AIxCC's design, the seven finalist architectures and results beyond the scoreboard. It reports that system stability and accuracy penalties decided rankings, that LLM-based systems found vulnerabilities a fuzzing baseline missed, and that among patches passing all automatic validation, manual review found semantic errors in 38-46% from baseline agents; the top two systems had 83.8% and 79.2% competition-scored patch accuracy.
Trail of Bits
Security firm; built the Buttercup cyber reasoning system for AIxCC.
DARPA reports that seven finalist cyber reasoning systems analyzed over 54 million lines of code, found 54 unique synthetic vulnerabilities in 63 challenges and patched 43, and found 18 real non-synthetic vulnerabilities with 11 patches. Team Atlanta won $4 million, Trail of Bits $3 million and Theori $1.5 million; DARPA and ARPA-H added $1.4 million for real-world integration and four systems were open-sourced on the day.
DARPA reports that in the AIxCC semifinal at DEF CON 32, nearly 40 cyber reasoning systems were tested on challenge projects based on Jenkins, the Linux kernel, Nginx, SQLite3 and Apache Tika. Competitors' systems found 22 unique synthetic vulnerabilities, patched 15, and found one real-world SQLite3 bug; seven teams advanced with $2 million each and must open-source their systems after the final.