Microsoft economists ran randomized controlled trials in which novices and security professionals completed incident summarization, script analysis, incident report and guided response tasks in a Defender XDR test environment, with half given Security Copilot. The January 2024 revision reports that novices with Copilot answered 35% more questions correctly and professionals were 7% more accurate, with both groups completing tasks faster.
Why it matters
It is one of the few controlled experiments measuring whether an LLM assistant changes SOC analyst performance, rather than relying on vendor anecdotes.
Key facts
As stated in the sources, with where to find them.
- Novice study: 149 subjects (tested October 2023); Copilot subjects got 35% more multiple-choice questions correct and were 26% faster holding accuracy constant.Whitepaper, 'Findings - novices'
- Professional study: 147 security professionals (tested December 2023 to January 2024); Copilot users were 7% more accurate on multiple-choice tasks and finished overall tasks 22% faster.Whitepaper, 'Findings - professionals'
- Authors note tasks were closely linked to Copilot's capabilities, which may overstate real-world gains.Whitepaper, novice findings discussion
Findings that cite this record
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Mar 24, 2025
Nov 17, 2025
Nov 22, 2024