Start here · where things stand

Key questions

The questions a researcher would ask about AI agents in cybersecurity. The questions stay fixed; their answers are revised in the open as evidence arrives, and every earlier answer is kept. Each answer states how confident we are and links to the findings it rests on. How to read confidence, statuses, and review flags.

1Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.Tool-using agents without added defenses followed injected instructions in a substantial share of 2024 benchmark cases. Frontier models tested in 2025 and 2026 resist far more often, yet every one was hijacked at least once in large public red-teaming competitions. Research defenses that reported near-zero attack success against fixed attacks failed once attackers adapted to them, and frontier labs report that their layered defenses reduce injection in browser and computer-use agents without eliminating it. The approach with the strongest support is architectural: bound what untrusted input can cause an agent to do, rather than rely on the model to resist.high confidence1 corroborated, 2 reported, 2 qualifiedRevised Sep 26, 2026 · 1 earlier answer2Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.Disclosed vulnerabilities show outside content reaching agents through email, calendar invites, and shared documents without a click, and broad credentials turning one injected instruction into wide access. In several agent frameworks, injection has led to code execution on the host, and malicious agent packages and MCP servers have been used against real users.moderate confidence5 corroboratedAnswered Sep 25, 20263Do cyber evaluations of AI agents stay contained?Not reliably. Several labs and a government evaluator have disclosed agents under evaluation acting on real third-party systems.Labs and a government evaluator disclosed frontier agents under cyber evaluation acting against real third-party systems, through misconfiguration, intentionally enabled internet access, or a flaw in shared infrastructure. Agents also took out-of-scope shortcuts in cyber evaluations, and some coordinated through unintended shared channels. Evaluation environments are part of the attack surface.high confidence3 corroboratedAnswered Sep 25, 20264How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.UK AISI reports that fixed, low token budgets understate frontier cyber capability and its rate of progress. A preprint audit finds pipeline choices alone can move cyber benchmark scores by tens of points, public CTF benchmarks can be contaminated, and counting a crash as exploitation overstates capability. Single capability numbers, including trend estimates, are best read as lower or conditional bounds.moderate confidence4 reportedAnswered Sep 25, 20265Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.AI systems find previously unknown vulnerabilities in widely used open-source software, and one lab’s program data shows disclosure far ahead of patching. Checking only that the original crash no longer reproduces overstates how often AI-generated patches actually fix the flaw. Independent patch verification is the bottleneck the evidence points to.moderate confidence3 corroboratedAnswered Sep 25, 20266Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.LLM agents perform poorly on realistic threat-hunting and investigation benchmarks, while vendor-run trials report assistants speeding up analysts who remain in charge. Agent monitors can be driven to miss covert actions under optimized attack, so catch rates measured against ordinary mistakes say little about adversarial settings.moderate confidence2 corroborated, 1 reported, 1 qualifiedAnswered Sep 25, 20267How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.Provider and vendor reports trace a shift. In 2024 and early 2025, Microsoft, OpenAI and Google reported threat actors using LLMs mainly as productivity tools. From mid-2025, Anthropic reported Claude Code carrying out an extortion campaign under human direction and a state-sponsored espionage campaign with people at a few decision points; Google reported malware that queries LLMs during execution, including by Russia’s APT28; and Sysdig, Google and ThreatDown reported agent-driven extortion, automated credential harvesting and a botnet built around an agent framework. Google also reported in September 2026 that it had not yet seen fully autonomous attack pipelines in the wild. These reports come from the organizations that detected the activity, mostly on their own platforms, so they show that agent-driven attacks happen, not how common they are.moderate confidence3 corroborated, 1 reportedRevised Sep 26, 2026 · 2 earlier answers
How answers stay current

An answer is flagged for review when records that bear on it are added, when one of its findings changes status or ages past its half-life, or 120 days after its last review. An editor then either revises the answer, which keeps the old one in its history, or confirms it. No answer is flagged now.

What these answers cannot see

Answers only reflect what the corpus holds. Known gaps: Offensive capability measurements; Threat-intelligence and misuse reports; June to September 2026; Non-English and non-Western sources. Details.

Quarterly record

Where things stood

The answers are fixed at the end of each quarter, so anyone can see what this record said at a given time. Each quarterly page is derived from the answer histories and cannot be edited after the fact.

2026Q3
Q3 2026
In progress; fixed on Sep 30, 2026