Scope: what this does not show
One benchmark with an LLM judge, run with lab safeguards disabled; the top model (Claude Mythos Preview) was unreleased, and most other models fell to zero with mitigations on.
Qualified: Still standing, but later work narrows how far it applies.
Evidence
May 11, 2026
ExploitGym benchmark measures whether AI agents can turn real vulnerabilities into working exploits
Table 3: Mythos Preview 157 and GPT-5.5 120 of 898. Table 5: with mitigations re-enabled, 45 and 21 remained.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Counting a crash as exploitation success overstates capability; most public models in May 2026 stalled before code execution on a browser engine. qualifies this findingOn hardened V8 targets, publicly deployed models rarely got past primitives to code execution; code execution at scale came only from an unreleased model.
Status history
- 2026-05-11ReportedExploitGym results. · record
- 2026-09-08QualifiedLater work shows cyber benchmark scores depend heavily on pipeline choices. · record
- 2026-09-25QualifiedcorrectionThe pipeline audit covered eight knowledge and multiple-choice benchmarks, not ExploitGym. The qualification rests on ExploitBench, where no publicly deployed model reached code execution on V8. · record