ExCyTIn-Bench builds threat-investigation questions from graphs of security logs collected in a controlled Azure tenant with simulated multi-step attacks, and asks agents to query the logs to answer them. In the July 2025 version the best model (o4-mini) reached a reward of 0.368; in the May 2026 revision, accepted at ICML 2026, the best (Claude Opus 4.5) reached 0.606, which the authors say leaves substantial headroom.
Why it matters
It is an open benchmark for the investigative, log-querying work of SOC analysts rather than multiple-choice knowledge.
Key facts
As stated in the sources, with where to find them.
- v1 (July 2025): 8 simulated multi-step attacks in a controlled Azure tenant, 57 log tables from Microsoft Sentinel and related services, 589 generated questions; average reward 0.249 across evaluated models, best 0.368 (o4-mini).arXiv v1 abstract; Table 2
- v3 (May 2026, ICML 2026 version): 7,542 generated questions from the same 57 log tables; best reward 0.606 (Claude Opus 4.5).arXiv v3 abstract; Section 1; Table 2
Findings that cite this record
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Mar 13, 2026
Apr 21, 2026
May 21, 2025
Aug 5, 2025
Sep 24, 2025
Jul 30, 2026