Topics/Defense

Autonomous defense

Agents that detect, contain, and respond without waiting for a person.

19 records2 findings2 openings3 benchmarks and toolsLatest record
Start here

Defenders' agents

What AI agents can and cannot yet do for security operations, from assisted analysts to autonomous response in simulation.

  1. Microsoft randomized controlled trial measures Security Copilot effect on analyst speed and accuracy
    An early controlled trial of an analyst assistant.
  2. TTCP releases CAGE Challenge 4, a multi-agent autonomous cyber defence environment
    A shared simulation for autonomous defense.
  3. UC Santa Cruz study integrates LLM agents into CAGE 4 and finds RL defenders still outperform them
    LLM defenders compared with reinforcement learning.
  4. LLM agents fall well short of reliable performance on realistic threat-investigation and threat-hunting benchmarks built from security logs.
    How agents fare on realistic investigation benchmarks.
  5. Year-long SOC fieldwork finds analysts reused an agentic AI companion's output in over 90% of tickets
    What happens in a real SOC.
  6. Why don't defensive agents get better with more compute?
    The open question.
RangeLanes
14 of 19 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
202420252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

14 records · newest first
Sep 2026
Sep 24, 2026
Google's PageBreak agent finds over 500 XSS bugs in its own web apps using deterministic validators
DefenseTool releaseGoogle

Google's Product Security team describes PageBreak, an internal agent mostly using Gemini models that hunts vulnerabilities in Google's first-party web applications and only reports findings confirmed by non-AI validators against running applications. Google reports over 500 XSS vulnerabilities found with near-zero false positives, while apps on its high-assurance web frameworks yielded only 2 XSS bugs as of 4 September 2026.

Mar 2026
Jan 2026
Jan 8, 2026
PNNL uses a Claude-based agent to speed adversary emulation against a water treatment plant model
DefensePaperAnthropic, Pacific Northwest National Laboratory, CISA

Anthropic reports that Pacific Northwest National Laboratory built a scaffold around Claude Sonnet 4 to automate adversary emulation against a high-fidelity cyber-physical model of a water treatment plant used for CISA. PNNL estimates attack reconstruction took three hours instead of multiple weeks; in one run the model switched to a different known privilege-escalation technique when a provided tool failed.

Dec 2025
Dec 16, 2025
NIST releases preliminary draft Cyber AI Profile (IR 8596) under CSF 2.0
PolicyFrameworkNIST

NIST published a preliminary draft Cybersecurity Framework Profile for Artificial Intelligence, aligned with CSF 2.0. It is organized around three focus areas: securing AI systems, using AI for cyber defense, and thwarting AI-enabled cyberattacks, with comments due January 30, 2026.

Dec 3, 2025
CISA, ASD and partners issue principles for securely integrating AI, including agents, into OT
PolicyGuidanceCISA, Australian Signals Directorate (ACSC), NSA Artificial Intelligence Security Center

CISA and the Australian Signals Directorate, with NSA, FBI and national cyber agencies of Canada, Germany, the Netherlands, New Zealand and the UK, published four principles for integrating AI into operational technology. The guidance explicitly covers machine learning, LLM-based AI and AI agents because of the security and safety challenges they pose in industrial environments.

Nov 2025
Nov 17, 2025
Microsoft RCT finds phishing triage agent raised analysts' true positives per minute up to 6.5x
DefensePaperMicrosoft

Microsoft reports a randomized controlled trial of its own Security Copilot Phishing Triage Agent. In the trial, 167 external security analysts each triaged a 25-email queue drawn from a curated corpus of emails reported by Microsoft employees. In the scenario where the agent classified every corpus email correctly, analysts with the agent found 6.5 times as many true positives per minute as the control group and scored 77% higher on F1; with the agent's accuracy set to 80% and a 20% malicious rate, the productivity gain fell to 3.1 times. Analysts with the agent spent 53% more time on malicious emails and did not simply confirm its malicious verdicts, but they were more likely to accept its benign verdicts, including planted false negatives.

Sep 2025
Aug 2025
Aug 8, 2025
AIxCC final: Team Atlanta wins as systems patch 43 of 54 found synthetic bugs and find 18 real ones
DefenseCompetitionDARPA, ARPA-H, Team Atlanta

DARPA reports that seven finalist cyber reasoning systems analyzed over 54 million lines of code, found 54 unique synthetic vulnerabilities in 63 challenges and patched 43, and found 18 real non-synthetic vulnerabilities with 11 patches. Team Atlanta won $4 million, Trail of Bits $3 million and Theori $1.5 million; DARPA and ARPA-H added $1.4 million for real-world integration and four systems were open-sourced on the day.

Aug 5, 2025
Microsoft's Project Ire agent autonomously reverse engineers and classifies malware
DefenseTool releaseMicrosoft

Microsoft Research describes Project Ire, a prototype LLM agent that uses decompilers and binary analysis tools to reverse engineer software and classify it as malicious or benign, producing an auditable chain-of-evidence report. Microsoft reports 0.98 precision and 0.83 recall on a Windows driver dataset, but 0.26 recall on about 4,000 hard real-world files, and plans to deploy it in Defender as Binary Analyzer.

Jun 2025
Jun 27, 2025
Paper proposes test and evaluation process with effectiveness metrics for RL cyber defence agents
DefensePaper

A paper in Applied AI Letters by QinetiQ researchers sets out a test and evaluation process for cyber defence agents covering performance, effectiveness, resilience and generalisability, and demonstrates its low-fidelity stage on CAGE Challenge 2 RL agents in CybORG. It introduces Measures of Effectiveness tailored to cyber defence alongside RL reward and tests agents under environment perturbations not seen in training.

May 2025
May 7, 2025
UC Santa Cruz study integrates LLM agents into CAGE 4 and finds RL defenders still outperform them
DefensePaperUC Santa Cruz

Researchers led by UC Santa Cruz integrated LLM agents into the CybORG CAGE 4 multi-agent defence environment and proposed a communication protocol for mixed LLM and RL teams. In their runs an all-RL team scored far better reward than an all-LLM (GPT-4o-mini) team and acted about 104 times faster, though the authors highlight LLM explainability and note the environment was designed for RL agents.

Mar 2025
Mar 24, 2025
Microsoft announces Security Copilot agents for phishing triage, alert triage and remediation
DefenseTool releaseMicrosoft, OneTrust, Aviatrix

Microsoft announced Microsoft-built Security Copilot agents, including a Phishing Triage Agent in Defender, alert triage agents in Purview, a Conditional Access Optimization Agent, a Vulnerability Remediation Agent in Intune and a Threat Intelligence Briefing Agent, plus five partner agents. Preview was planned from April 2025; the announcement contains no evaluation of agent accuracy.

Nov 2024
Nov 22, 2024
Frontier Model Forum issue brief maps defensive uses of frontier AI in cybersecurity
PolicyGuidanceFrontier Model Forum

The Frontier Model Forum, an industry body of frontier labs, published an issue brief on using frontier AI for cyber defense. It lists use cases including process automation for incident response, natural-language querying and analysis, vulnerability discovery and fixing, open-source intelligence and training, and recommends designing for human-AI collaboration rather than full automation.

Findings

Research openings

Benchmarks and tools