Organizations/academic

Stanford University

1 records1 defense
May 21, 2025
BountyBench measures AI agents on detect, exploit and patch tasks from real bug bounties
DefenseBenchmarkStanford University, UC Berkeley

BountyBench, from Stanford-led researchers, builds 40 bug bounties across 25 real-world systems into 120 Detect, Exploit and Patch tasks with dollar values attached. In the first version the best Detect score was 5%, while OpenAI Codex CLI and Claude Code scored 90% and 87.5% on Patch, well above their Exploit scores. A July 2025 revision with more agents reported Codex CLI with o3-high at 12.5% on Detect and 90% on Patch.