CTI-REALM

Benchmark testing whether agents turn threat intelligence into validated detections.

Records citing CTI-REALM

Mar 13, 2026
Microsoft's CTI-REALM benchmark tests agents turning threat intel into validated detection rules
DefenseBenchmarkMicrosoft

CTI-REALM places agents in a tool-rich environment where they read threat intelligence reports, explore telemetry, iterate KQL queries and produce Sigma and KQL detection rules across Linux, AKS and Azure cloud scenarios. The paper's evaluation of 16 model configurations found Claude Opus 4.6 (High) best at 0.637, with cloud detection hardest; Microsoft's blog later added an early Claude Mythos Preview snapshot scoring 0.685.