Scope: what this does not show
Vendor-run trials of vendor products; assisted work, not autonomous response.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- LLM agents fall well short of reliable performance on realistic threat-investigation and threat-hunting benchmarks built from security logs. qualifies this findingGains appear in assisted work; autonomous performance on realistic tasks remains weak.
Key questions that rely on this finding
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Status history
- 2023-12-05ReportedFirst Security Copilot trial. · record