Chronicle/Defense & research

Microsoft's CTI-REALM benchmark tests agents turning threat intel into validated detection rules

DefenseBenchmarkSignificance assistant-drafted

CTI-REALM places agents in a tool-rich environment where they read threat intelligence reports, explore telemetry, iterate KQL queries and produce Sigma and KQL detection rules across Linux, AKS and Azure cloud scenarios. The paper's evaluation of 16 model configurations found Claude Opus 4.6 (High) best at 0.637, with cloud detection hardest; Microsoft's blog later added an early Claude Mythos Preview snapshot scoring 0.685.

Why it matters

It measures an end-to-end detection engineering workflow, a core SOC task that most security benchmarks skip.

Key facts

As stated in the sources, with where to find them.

  • 37 curated CTI reports; CTI-REALM-50 has 50 tasks across Linux, AKS and Azure cloud.Microsoft blog
  • Across 16 frontier model configurations, Claude Opus 4.6 (High) achieved the highest reward (0.637), followed by Claude Opus 4.5 (0.624) and the GPT-5 family.arXiv abstract
  • Scores fall from Linux (0.585) to AKS (0.517) to cloud (0.282); removing CTI-specific tools cut performance by up to 0.150.Microsoft blog, findings list
  • The blog, updated after publication, reports an early Claude Mythos Preview snapshot at 0.685.Microsoft blog, results

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records