Findings/soc-agents-weak-on-realistic-benchmarks

LLM agents fall well short of reliable performance on realistic threat-investigation and threat-hunting benchmarks built from security logs.

Corroboratedmeasured3 evidence records from 3 independent sourcesassistant-drafted
Scope: what this does not show

Benchmarks, several built by security vendors (Microsoft, Simbian); not measurements of deployed systems. CyberSOCEval is multiple-choice question answering, not an agentic task.

Corroborated: Supported by at least two independent sources.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2025-07-14ReportedExCyTIn-Bench results. · record
  2. 2025-09-24CorroboratedCyberSOCEval finds similar limits. · record
  3. 2026-09-25CorroboratedcorrectionCyberSOCEval tests multiple-choice question answering, not agents on investigation or hunting. Corroboration rests on Simbian's Cyber Defense Benchmark, where the best of five models flagged 3.8% of malicious events in raw logs. · record