Chronicle/Defense & research

Meta and CrowdStrike release CyberSOCEval benchmarks for malware analysis and threat intel reasoning

DefenseBenchmarkSignificance assistant-drafted

CyberSOCEval adds two open-source SOC benchmarks to CyberSecEval 4: malware analysis questions built from sandbox detonation reports, and threat intelligence reasoning over unstructured reports. The authors find larger, newer models do better, reasoning models gain less than in coding and math, and current models are far from saturating the tasks.

Why it matters

It gives defenders an open benchmark grounded in real sandbox and threat-report data rather than generic security trivia.

Key facts

As stated in the sources, with where to find them.

  • CyberSOCEval covers two tasks, Malware Analysis and Threat Intelligence Reasoning, within CyberSecEval 4.Abstract
  • Reasoning models using test-time scaling do not get the boost seen in coding and math; models are far from saturating the benchmark.Abstract

Findings that cite this record

Key questions this bears on

Sources

Related records