Chronicle/Defense & research

Cost-aware evaluation finds defensive SOC agents do not scale with compute like offensive CTF agents

DefensePaperSignificance assistant-drafted

Researchers evaluate security agents at fixed cost levels on offensive Cybench challenges and defensive Splunk BOTS v1 investigations, splitting spend into inference and tool use. They find offensive success rises with test-time compute, while defensive investigation depends more on disciplined tool use and telemetry navigation, and argue benchmarks should report cost and operational fit alongside success.

Why it matters

It argues that security-agent benchmarks reporting only peak success under generous budgets miss cost and operational fit, and that defensive SOC work does not reward extra compute the way offensive CTFs do.

Key facts

As stated in the sources, with where to find them.

  • Evaluation uses Cybench (offensive) and Splunk BOTS v1 (defensive), comparing models at fixed cost levels and splitting spend into inference and tool costs.Abstract
  • Scaled open-weight models can approach frontier proprietary systems on offensive CTFs while remaining cost-competitive.Abstract
  • The defensive analysis covers only the 31 scored BOTS v1 questions; the authors present the SOC finding as an evaluation-design result, not a verdict on production SOC readiness.Section 7, Limitations

Findings that cite this record

Key questions this bears on

Sources

Related records