ExCyTIn-Bench

Benchmark of multi-step threat investigations over security logs.

Records citing ExCyTIn-Bench

Jul 14, 2025
Microsoft's ExCyTIn-Bench evaluates LLM agents on multi-step threat investigation over Sentinel logs
DefenseBenchmarkMicrosoft

ExCyTIn-Bench builds threat-investigation questions from graphs of security logs collected in a controlled Azure tenant with simulated multi-step attacks, and asks agents to query the logs to answer them. In the July 2025 version the best model (o4-mini) reached a reward of 0.368; in the May 2026 revision, accepted at ICML 2026, the best (Claude Opus 4.5) reached 0.606, which the authors say leaves substantial headroom.