CyberSOCEval adds two open-source SOC benchmarks to CyberSecEval 4: malware analysis questions built from sandbox detonation reports, and threat intelligence reasoning over unstructured reports. The authors find larger, newer models do better, reasoning models gain less than in coding and math, and current models are far from saturating the tasks.
Why it matters
It gives defenders an open benchmark grounded in real sandbox and threat-report data rather than generic security trivia.
Key facts
As stated in the sources, with where to find them.
- CyberSOCEval covers two tasks, Malware Analysis and Threat Intelligence Reasoning, within CyberSecEval 4.Abstract
- Reasoning models using test-time scaling do not get the boost seen in coding and math; models are far from saturating the benchmark.Abstract
Findings that cite this record
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Mar 13, 2026
Jul 14, 2025
Jul 16, 2026
Apr 21, 2026
Mar 30, 2026
Apr 29, 2025