BountyBench

Benchmark built from real bug-bounty cases with detect, exploit, and patch tasks.

Records citing BountyBench

May 21, 2025
BountyBench measures AI agents on detect, exploit and patch tasks from real bug bounties
DefenseBenchmarkStanford University, UC Berkeley

BountyBench, from Stanford-led researchers, builds 40 bug bounties across 25 real-world systems into 120 Detect, Exploit and Patch tasks with dollar values attached. In the first version the best Detect score was 5%, while OpenAI Codex CLI and Claude Code scored 90% and 87.5% on Patch, well above their Exploit scores. A July 2025 revision with more agents reported Codex CLI with o3-high at 12.5% on Detect and 90% on Patch.