Carnegie Mellon researchers released ExploitBench, which scores exploitation progress on 41 V8 JavaScript-engine vulnerabilities across 16 flags from reaching the bug through arbitrary read/write, control-flow hijack and code execution. The paper reports that public models routinely reach and crash vulnerable code but rarely achieve arbitrary code execution, while one private frontier model succeeded on roughly half of cases.
Why it matters
Graded scoring separates reaching or crashing a bug from building a working exploit, which crash-as-success benchmarks conflate.
Key facts
As stated in the sources, with where to find them.
- 41 V8 vulnerabilities; 16 measurable capability flags; 8 public frontier models plus 1 private model evaluated.Abstract
- The private frontier model reached arbitrary code execution on approximately half of cases.Abstract
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.
Sources
Related records
Jul 23, 2026
Jul 2, 2026
May 29, 2026
May 13, 2026
May 12, 2026
May 11, 2026