Findings/frontier-models-produce-working-exploits

On ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.

Qualifiedmeasured1 evidence record from 1 independent sourceassistant-drafted
Scope: what this does not show

One benchmark with an LLM judge, run with lab safeguards disabled; the top model (Claude Mythos Preview) was unreleased, and most other models fell to zero with mitigations on.

Qualified: Still standing, but later work narrows how far it applies.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Status history

  1. 2026-05-11ReportedExploitGym results. · record
  2. 2026-09-08QualifiedLater work shows cyber benchmark scores depend heavily on pipeline choices. · record
  3. 2026-09-25QualifiedcorrectionThe pipeline audit covered eight knowledge and multiple-choice benchmarks, not ExploitGym. The qualification rests on ExploitBench, where no publicly deployed model reached code execution on V8. · record