Scope: what this does not show
One benchmark on one target class.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
May 13, 2026
ExploitBench grades AI exploit development as a 16-step capability ladder on V8 bugs
Of 41 V8 bugs, crashes were common but no publicly deployed model reached arbitrary code execution in the primary arm; the unreleased Mythos Preview did on 18.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies On ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.On hardened V8 targets, publicly deployed models rarely got past primitives to code execution; code execution at scale came only from an unreleased model.
Key questions that rely on this finding
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Status history
- 2026-05-13ReportedExploitBench separates crashes from exploitation progress. · record