Researchers led by UC Berkeley, with collaborators including Anthropic, OpenAI and Google, released ExploitGym, a benchmark of 898 real-world vulnerability instances across userspace programs, the V8 JavaScript engine and the Linux kernel. Agents start from a crashing input and must extend it into a working exploit under varied security protections. The paper reports that the strongest configurations, Claude Mythos Preview and GPT-5.5, produced working exploits for 157 and 120 instances respectively.
Why it matters
ExploitGym became a shared exploit-development yardstick in 2026 lab system cards and was the evaluation running during the Hugging Face intrusion.
Key facts
As stated in the sources, with where to find them.
- 898 instances from real-world vulnerabilities in three domains: userspace programs, V8, and the Linux kernel.Abstract
- Claude Mythos Preview produced working exploits for 157 instances and GPT-5.5 for 120.Abstract
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.
Sources
Related records
May 13, 2026
Jul 21, 2026
Mar 13, 2026
Aug 4, 2026
May 29, 2026
Mar 1, 2026