Researchers evaluate security agents at fixed cost levels on offensive Cybench challenges and defensive Splunk BOTS v1 investigations, splitting spend into inference and tool use. They find offensive success rises with test-time compute, while defensive investigation depends more on disciplined tool use and telemetry navigation, and argue benchmarks should report cost and operational fit alongside success.
Why it matters
It argues that security-agent benchmarks reporting only peak success under generous budgets miss cost and operational fit, and that defensive SOC work does not reward extra compute the way offensive CTFs do.
Key facts
As stated in the sources, with where to find them.
- Evaluation uses Cybench (offensive) and Splunk BOTS v1 (defensive), comparing models at fixed cost levels and splitting spend into inference and tool costs.Abstract
- Scaled open-weight models can approach frontier proprietary systems on offensive CTFs while remaining cost-competitive.Abstract
- The defensive analysis covers only the 31 scored BOTS v1 questions; the authors present the SOC finding as an evaluation-design result, not a verdict on production SOC readiness.Section 7, Limitations
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Apr 21, 2026
Sep 8, 2026
Jul 30, 2026
Jul 21, 2026
Jul 2, 2026
May 29, 2026