CTI-REALM places agents in a tool-rich environment where they read threat intelligence reports, explore telemetry, iterate KQL queries and produce Sigma and KQL detection rules across Linux, AKS and Azure cloud scenarios. The paper's evaluation of 16 model configurations found Claude Opus 4.6 (High) best at 0.637, with cloud detection hardest; Microsoft's blog later added an early Claude Mythos Preview snapshot scoring 0.685.
Why it matters
It measures an end-to-end detection engineering workflow, a core SOC task that most security benchmarks skip.
Key facts
As stated in the sources, with where to find them.
- 37 curated CTI reports; CTI-REALM-50 has 50 tasks across Linux, AKS and Azure cloud.Microsoft blog
- Across 16 frontier model configurations, Claude Opus 4.6 (High) achieved the highest reward (0.637), followed by Claude Opus 4.5 (0.624) and the GPT-5 family.arXiv abstract
- Scores fall from Linux (0.585) to AKS (0.517) to cloud (0.282); removing CTI-specific tools cut performance by up to 0.150.Microsoft blog, findings list
- The blog, updated after publication, reports an early Claude Mythos Preview snapshot at 0.685.Microsoft blog, results
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Mar 1, 2026
Apr 21, 2026
Jul 21, 2026
Jul 14, 2025
May 11, 2026
May 1, 2026