Chronicle/Defense & research

BountyBench measures AI agents on detect, exploit and patch tasks from real bug bounties

DefenseBenchmarkSignificance assistant-drafted

BountyBench, from Stanford-led researchers, builds 40 bug bounties across 25 real-world systems into 120 Detect, Exploit and Patch tasks with dollar values attached. In the first version the best Detect score was 5%, while OpenAI Codex CLI and Claude Code scored 90% and 87.5% on Patch, well above their Exploit scores. A July 2025 revision with more agents reported Codex CLI with o3-high at 12.5% on Detect and 90% on Patch.

Why it matters

It puts offensive and defensive agent performance on the same real codebases and expresses results in bounty dollars.

Key facts

As stated in the sources, with where to find them.

  • 25 systems, 40 bug bounties with awards from $10 to $30,485, covering 9 of the OWASP Top 10; 120 tasks.Abstract; Section 1
  • v1 (21 May 2025, 5 agents): best Detect 5% (Codex CLI $2,400; Claude Code $1,350); Codex CLI 90% Patch ($14,422); Claude 3.7 Sonnet Thinking custom agent 67.5% Exploit; Codex CLI and Claude Code Patch 90% and 87.5% vs Exploit 32.5% and 57.5%.v1 abstract
  • v2 (10 July 2025, 8 agents): Codex CLI o3-high 12.5% Detect ($3,720) and 90% Patch ($14,152); Codex CLI o4-mini 90% Patch ($14,422); Codex CLI o3-high, o4-mini and Claude Code scored 90%, 90% and 87.5% on Patch vs 47.5%, 32.5% and 57.5% on Exploit.v2 abstract

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records