Scope: what this does not show
Benchmarks, several built by security vendors (Microsoft, Simbian); not measurements of deployed systems. CyberSOCEval is multiple-choice question answering, not an agentic task.
Corroborated: Supported by at least two independent sources.
Evidence
Jul 14, 2025
Microsoft's ExCyTIn-Bench evaluates LLM agents on multi-step threat investigation over Sentinel logs
Best model reward 0.606 on log-investigation questions.
Sep 24, 2025
Meta and CrowdStrike release CyberSOCEval benchmarks for malware analysis and threat intel reasoning
Multiple-choice questions over malware sandbox and threat reports, not agentic tasks; models far from saturating.
Apr 21, 2026
Threat-hunting benchmark finds best LLM agent flags only 3.8% of malicious events in raw logs
The best agent flagged 3.8% of malicious events in raw logs.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies Controlled trials run by Microsoft report that its security assistants make analysts faster and more accurate.Gains appear in assisted work; autonomous performance on realistic tasks remains weak.
Key questions that rely on this finding
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Status history
- 2025-07-14ReportedExCyTIn-Bench results. · record
- 2025-09-24CorroboratedCyberSOCEval finds similar limits. · record
- 2026-09-25CorroboratedcorrectionCyberSOCEval tests multiple-choice question answering, not agents on investigation or hunting. Corroboration rests on Simbian's Cyber Defense Benchmark, where the best of five models flagged 3.8% of malicious events in raw logs. · record