Chronicle/Capability & gating

ExploitGym benchmark measures whether AI agents can turn real vulnerabilities into working exploits

CapabilityBenchmarkSignificance assistant-drafted

Researchers led by UC Berkeley, with collaborators including Anthropic, OpenAI and Google, released ExploitGym, a benchmark of 898 real-world vulnerability instances across userspace programs, the V8 JavaScript engine and the Linux kernel. Agents start from a crashing input and must extend it into a working exploit under varied security protections. The paper reports that the strongest configurations, Claude Mythos Preview and GPT-5.5, produced working exploits for 157 and 120 instances respectively.

Why it matters

ExploitGym became a shared exploit-development yardstick in 2026 lab system cards and was the evaluation running during the Hugging Face intrusion.

Key facts

As stated in the sources, with where to find them.

  • 898 instances from real-world vulnerabilities in three domains: userspace programs, V8, and the Linux kernel.Abstract
  • Claude Mythos Preview produced working exploits for 157 instances and GPT-5.5 for 120.Abstract

Findings that cite this record

Key questions this bears on

Sources

Related records