Chronicle/Defense & research

Microsoft's ExCyTIn-Bench evaluates LLM agents on multi-step threat investigation over Sentinel logs

DefenseBenchmarkSignificance assistant-drafted

ExCyTIn-Bench builds threat-investigation questions from graphs of security logs collected in a controlled Azure tenant with simulated multi-step attacks, and asks agents to query the logs to answer them. In the July 2025 version the best model (o4-mini) reached a reward of 0.368; in the May 2026 revision, accepted at ICML 2026, the best (Claude Opus 4.5) reached 0.606, which the authors say leaves substantial headroom.

Why it matters

It is an open benchmark for the investigative, log-querying work of SOC analysts rather than multiple-choice knowledge.

Key facts

As stated in the sources, with where to find them.

  • v1 (July 2025): 8 simulated multi-step attacks in a controlled Azure tenant, 57 log tables from Microsoft Sentinel and related services, 589 generated questions; average reward 0.249 across evaluated models, best 0.368 (o4-mini).arXiv v1 abstract; Table 2
  • v3 (May 2026, ICML 2026 version): 7,542 generated questions from the same 57 log tables; best reward 0.606 (Claude Opus 4.5).arXiv v3 abstract; Section 1; Table 2

Findings that cite this record

Key questions this bears on

Sources

Related records