BountyBench, from Stanford-led researchers, builds 40 bug bounties across 25 real-world systems into 120 Detect, Exploit and Patch tasks with dollar values attached. In the first version the best Detect score was 5%, while OpenAI Codex CLI and Claude Code scored 90% and 87.5% on Patch, well above their Exploit scores. A July 2025 revision with more agents reported Codex CLI with o3-high at 12.5% on Detect and 90% on Patch.
Why it matters
It puts offensive and defensive agent performance on the same real codebases and expresses results in bounty dollars.
Key facts
As stated in the sources, with where to find them.
- 25 systems, 40 bug bounties with awards from $10 to $30,485, covering 9 of the OWASP Top 10; 120 tasks.Abstract; Section 1
- v1 (21 May 2025, 5 agents): best Detect 5% (Codex CLI $2,400; Claude Code $1,350); Codex CLI 90% Patch ($14,422); Claude 3.7 Sonnet Thinking custom agent 67.5% Exploit; Codex CLI and Claude Code Patch 90% and 87.5% vs Exploit 32.5% and 57.5%.v1 abstract
- v2 (10 July 2025, 8 agents): Codex CLI o3-high 12.5% Detect ($3,720) and 90% Patch ($14,152); Codex CLI o4-mini 90% Patch ($14,422); Codex CLI o3-high, o4-mini and Claude Code scored 90%, 90% and 87.5% on Patch vs 47.5%, 32.5% and 57.5% on Exploit.v2 abstract
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
Sources
Related records
Jun 13, 2025
Oct 3, 2025
Jul 14, 2025
Aug 8, 2025
May 12, 2026
Feb 7, 2026