Cybench

Benchmark of professional-level capture-the-flag tasks for evaluating language-model agents on offensive security work.

Records citing Cybench

Jul 16, 2026
Cost-aware evaluation finds defensive SOC agents do not scale with compute like offensive CTF agents
DefensePaperarXiv

Researchers evaluate security agents at fixed cost levels on offensive Cybench challenges and defensive Splunk BOTS v1 investigations, splitting spend into inference and tool use. They find offensive success rises with test-time compute, while defensive investigation depends more on disciplined tool use and telemetry navigation, and argue benchmarks should report cost and operational fit alongside success.

May 21, 2026
Position paper argues agent security benchmarks suffer from hackable environments, staleness and runtime noise
DefensePaperarXiv

Abdelnabi, Hicks, Rieck and Sadeghi argue that security evaluations of agents face three problems: agents can break the benchmark environment instead of solving the task, static benchmarks such as CyberGym and Cybench age as vulnerabilities are patched or leak, and stochastic behavior, agent-written code and external dependencies make single runs unreliable. They propose stronger environment isolation, canary tokens to detect cheating, continually updated or live benchmarks, reporting worst-case results and variance, and benchmark introspection, which they call a holistic first step.

Oct 3, 2025
Anthropic says it trained Claude Sonnet 4.5 for defensive vulnerability finding and patching
DefensePaperAnthropic, HackerOne, CrowdStrike

Anthropic reports that a small team focused Claude Sonnet 4.5 training on finding and patching vulnerabilities and on testing simulated security infrastructure, while avoiding enhancements that clearly favour offence. It reports Sonnet 4.5 results on Cybench and CyberGym, a preliminary patching study in which 15% of patches were judged semantically equivalent to human references, and invites work on SOC and SIEM automation.