<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>SOC automation · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/soc-automation/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/soc-automation/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on soc automation, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>Correction to a finding (reconfirmed as corroborated): LLM agents fall well short of reliable performance on realistic threat-investigation and threat-hunting benchmarks built from security logs.</title>
<link>https://agentic-cyber-explorer.pages.dev/findings/soc-agents-weak-on-realistic-benchmarks/</link>
<guid isPermaLink="false">correction:soc-agents-weak-on-realistic-benchmarks:2026-09-25:corroborated</guid>
<pubDate>Fri, 25 Sep 2026 12:00:00 GMT</pubDate>
<category>Correction</category>
<description>Correction: CyberSOCEval tests multiple-choice question answering, not agents on investigation or hunting. Corroboration rests on Simbian's Cyber Defense Benchmark, where the best of five models flagged 3.8% of malicious events in raw logs.</description>
</item>
<item>
<title>Year-long SOC fieldwork finds analysts reused an agentic AI companion's output in over 90% of tickets</title>
<link>https://agentic-cyber-explorer.pages.dev/events/usf-soc-agentic-ai-companion-deployment-2026/</link>
<guid isPermaLink="false">event:usf-soc-agentic-ai-companion-deployment-2026</guid>
<pubDate>Sat, 05 Sep 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>University of South Florida researchers embedded in a working SOC for over a year built and deployed an LLM-based agentic companion to handle high-volume, low-priority tickets, with analysts using it in the final four months. They report that companion outputs were reused in analysts' closing reports in more than 90% of cases, and that analysts who shaped the companion's behaviour came to trust it more. It is field evidence from a real SOC, not a benchmark, on how analysts adopt and trust an AI triage agent.</description>
</item>
<item>
<title>Cost-aware evaluation finds defensive SOC agents do not scale with compute like offensive CTF agents</title>
<link>https://agentic-cyber-explorer.pages.dev/events/cost-aware-security-agent-evaluation-2026/</link>
<guid isPermaLink="false">event:cost-aware-security-agent-evaluation-2026</guid>
<pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Researchers evaluate security agents at fixed cost levels on offensive Cybench challenges and defensive Splunk BOTS v1 investigations, splitting spend into inference and tool use. They find offensive success rises with test-time compute, while defensive investigation depends more on disciplined tool use and telemetry navigation, and argue benchmarks should report cost and operational fit alongside success. It argues that security-agent benchmarks reporting only peak success under generous budgets miss cost and operational fit, and that defensive SOC work does not reward extra compute the way offensive CTFs do.</description>
</item>
<item>
<title>Threat-hunting benchmark finds best LLM agent flags only 3.8% of malicious events in raw logs</title>
<link>https://agentic-cyber-explorer.pages.dev/events/simbian-cyber-defense-benchmark-threat-hunting-2026/</link>
<guid isPermaLink="false">event:simbian-cyber-defense-benchmark-threat-hunting-2026</guid>
<pubDate>Tue, 21 Apr 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>A technical report from security vendor Simbian AI presents the Cyber Defense Benchmark, which asks agents to hunt through 75,000-135,000 raw Windows event log records per episode, with no guiding questions, and flag the timestamps of malicious events drawn from 106 OTRF attack procedures. In the first version, the best of five frontier models (Claude Opus 4.6) flagged only 3.8% of malicious events on average and no model met the authors' bar of 50% recall on every ATT&amp;CK tactic. A revision two days later, with more models and a new coverage metric, reached the same no-pass conclusion. It contrasts strong LLM scores on curated security Q&amp;A with very weak performance on open-ended threat hunting.</description>
</item>
<item>
<title>Microsoft's CTI-REALM benchmark tests agents turning threat intel into validated detection rules</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-cti-realm-benchmark-2026/</link>
<guid isPermaLink="false">event:microsoft-cti-realm-benchmark-2026</guid>
<pubDate>Fri, 13 Mar 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>CTI-REALM places agents in a tool-rich environment where they read threat intelligence reports, explore telemetry, iterate KQL queries and produce Sigma and KQL detection rules across Linux, AKS and Azure cloud scenarios. The paper's evaluation of 16 model configurations found Claude Opus 4.6 (High) best at 0.637, with cloud detection hardest; Microsoft's blog later added an early Claude Mythos Preview snapshot scoring 0.685. It measures an end-to-end detection engineering workflow, a core SOC task that most security benchmarks skip.</description>
</item>
<item>
<title>Microsoft RCT finds phishing triage agent raised analysts' true positives per minute up to 6.5x</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-phishing-triage-agent-rct-2025/</link>
<guid isPermaLink="false">event:microsoft-phishing-triage-agent-rct-2025</guid>
<pubDate>Mon, 17 Nov 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Microsoft reports a randomized controlled trial of its own Security Copilot Phishing Triage Agent. In the trial, 167 external security analysts each triaged a 25-email queue drawn from a curated corpus of emails reported by Microsoft employees. In the scenario where the agent classified every corpus email correctly, analysts with the agent found 6.5 times as many true positives per minute as the control group and scored 77% higher on F1; with the agent's accuracy set to 80% and a 20% malicious rate, the productivity gain fell to 3.1 times. Analysts with the agent spent 53% more time on malicious emails and did not simply confirm its malicious verdicts, but they were more likely to accept its benign verdicts, including planted false negatives. It is one of the few randomized measurements of a commercial SOC triage agent's effect on analysts, and it reports automation bias toward the agent's benign verdicts alongside the productivity gains. It is a vendor study of its own product in a controlled task, not live operations.</description>
</item>
<item>
<title>Meta and CrowdStrike release CyberSOCEval benchmarks for malware analysis and threat intel reasoning</title>
<link>https://agentic-cyber-explorer.pages.dev/events/meta-crowdstrike-cybersoceval-2025/</link>
<guid isPermaLink="false">event:meta-crowdstrike-cybersoceval-2025</guid>
<pubDate>Wed, 24 Sep 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>CyberSOCEval adds two open-source SOC benchmarks to CyberSecEval 4: malware analysis questions built from sandbox detonation reports, and threat intelligence reasoning over unstructured reports. The authors find larger, newer models do better, reasoning models gain less than in coding and math, and current models are far from saturating the tasks. It gives defenders an open benchmark grounded in real sandbox and threat-report data rather than generic security trivia.</description>
</item>
<item>
<title>Microsoft's Project Ire agent autonomously reverse engineers and classifies malware</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-project-ire-malware-classification-2025/</link>
<guid isPermaLink="false">event:microsoft-project-ire-malware-classification-2025</guid>
<pubDate>Tue, 05 Aug 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Microsoft Research describes Project Ire, a prototype LLM agent that uses decompilers and binary analysis tools to reverse engineer software and classify it as malicious or benign, producing an auditable chain-of-evidence report. Microsoft reports 0.98 precision and 0.83 recall on a Windows driver dataset, but 0.26 recall on about 4,000 hard real-world files, and plans to deploy it in Defender as Binary Analyzer. It is a rare defensive-agent announcement that publishes both strong and weak results, including low recall on hard samples.</description>
</item>
<item>
<title>Microsoft's ExCyTIn-Bench evaluates LLM agents on multi-step threat investigation over Sentinel logs</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-excytin-bench-2025/</link>
<guid isPermaLink="false">event:microsoft-excytin-bench-2025</guid>
<pubDate>Mon, 14 Jul 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>ExCyTIn-Bench builds threat-investigation questions from graphs of security logs collected in a controlled Azure tenant with simulated multi-step attacks, and asks agents to query the logs to answer them. In the July 2025 version the best model (o4-mini) reached a reward of 0.368; in the May 2026 revision, accepted at ICML 2026, the best (Claude Opus 4.5) reached 0.606, which the authors say leaves substantial headroom. It is an open benchmark for the investigative, log-querying work of SOC analysts rather than multiple-choice knowledge.</description>
</item>
<item>
<title>Google announces Sec-Gemini v1, an experimental model for security operations workflows</title>
<link>https://agentic-cyber-explorer.pages.dev/events/google-sec-gemini-v1-2025/</link>
<guid isPermaLink="false">event:google-sec-gemini-v1-2025</guid>
<pubDate>Fri, 04 Apr 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Google announced Sec-Gemini v1, an experimental model combining Gemini with Google Threat Intelligence, OSV and Mandiant data for tasks such as incident root-cause analysis and vulnerability impact assessment. Google reports it outperforms other models by at least 11% on CTI-MCQ and 10.5% on CTI-Root Cause Mapping, and offered free research access to selected organizations. It is an example of a defender-specialized model whose advantage comes from integrated threat-intelligence data rather than only model scale.</description>
</item>
<item>
<title>Microsoft announces Security Copilot agents for phishing triage, alert triage and remediation</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-security-copilot-agents-2025/</link>
<guid isPermaLink="false">event:microsoft-security-copilot-agents-2025</guid>
<pubDate>Mon, 24 Mar 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Microsoft announced Microsoft-built Security Copilot agents, including a Phishing Triage Agent in Defender, alert triage agents in Purview, a Conditional Access Optimization Agent, a Vulnerability Remediation Agent in Intune and a Threat Intelligence Briefing Agent, plus five partner agents. Preview was planned from April 2025; the announcement contains no evaluation of agent accuracy. It marked a major vendor's shift from assistant-style copilots to semi-autonomous triage and remediation agents inside SOC tooling.</description>
</item>
<item>
<title>Frontier Model Forum issue brief maps defensive uses of frontier AI in cybersecurity</title>
<link>https://agentic-cyber-explorer.pages.dev/events/fmf-issue-brief-ai-for-cyber-defense-2024/</link>
<guid isPermaLink="false">event:fmf-issue-brief-ai-for-cyber-defense-2024</guid>
<pubDate>Fri, 22 Nov 2024 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The Frontier Model Forum, an industry body of frontier labs, published an issue brief on using frontier AI for cyber defense. It lists use cases including process automation for incident response, natural-language querying and analysis, vulnerability discovery and fixing, open-source intelligence and training, and recommends designing for human-AI collaboration rather than full automation. It records the lab consortium's stated position on defensive agent use before autonomous defense became a policy priority in 2026.</description>
</item>
<item>
<title>Microsoft randomized controlled trial measures Security Copilot effect on analyst speed and accuracy</title>
<link>https://agentic-cyber-explorer.pages.dev/events/microsoft-security-copilot-rct-2023/</link>
<guid isPermaLink="false">event:microsoft-security-copilot-rct-2023</guid>
<pubDate>Tue, 05 Dec 2023 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Microsoft economists ran randomized controlled trials in which novices and security professionals completed incident summarization, script analysis, incident report and guided response tasks in a Defender XDR test environment, with half given Security Copilot. The January 2024 revision reports that novices with Copilot answered 35% more questions correctly and professionals were 7% more accurate, with both groups completing tasks faster. It is one of the few controlled experiments measuring whether an LLM assistant changes SOC analyst performance, rather than relying on vendor anecdotes.</description>
</item>
</channel>
</rss>
