Changelog
Every change to the record, newest first: records added, findings recorded or moved on new evidence, key questions answered or revised, and our own corrections. Nothing is edited silently.
When new evidence changes a finding, it appears here as Finding moved and on the Desk. When we find that our own earlier judgment was wrong, the fix is appended to the finding’s history as a Correction, listed only here, and never shown as a change in the field.
September 2026
- Sep 26, 2026Key question revised
Can prompt injection against AI agents be reliably defended? was revised: Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach. Revised because newer competitions show much lower injection success on current frontier models, which qualifies the 2024 benchmark finding. The conclusion is unchanged: no model or defense has eliminated injection.
- Sep 26, 2026Key question revised
How are attackers using AI agents in real operations? was revised: Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs. Correction: the extortion campaign ran under human direction, security vendors are among the sources, and Google has not yet seen fully autonomous pipelines in the wild. Government threat reports are still missing from the corpus.
- Sep 26, 2026Finding movedUndefended tool-using agents follow injected instructions in a substantial share of benchmark cases.Public red-teaming competitions on 2025 and 2026 frontier models with built-in safeguards report much lower per-model success (0.5% to 8.5% in 2026), though every model was hijacked at least once. The substantial rates describe 2024 models and benchmarks.
- Sep 25, 2026Key question revised
How are attackers using AI agents in real operations? was revised: Increasingly to run parts of intrusions: providers report agent-driven espionage, extortion and credential theft, and malware that queries LLMs as it runs. Revised after twelve threat-intelligence and malware reports from 2024 to September 2026 (Anthropic, ESET, Google, Microsoft with OpenAI, Sysdig and ThreatDown) were added, closing most of the coverage gap the first answer described.
- Sep 25, 2026CorrectionAttackers who adapt to a defense defeat most published prompt-injection defenses that reported near-zero success against static attacks.The 2025 US AISI and Google DeepMind entries did not test published defenses with near-zero reported success, and DeepMind shares authors with the primary study. 'The Attacker Moves Second' is the primary evidence; no independent replication is recorded yet.
- Sep 25, 2026CorrectionIn DARPA's AI Cyber Challenge, autonomous systems patched most of the synthetic vulnerabilities they found.The 16-21% figure is competition-scored submission accuracy and does not reduce DARPA's 43 counted patches. The qualification now rests on PatchBench: agents from top AIxCC teams lose much of their solve rate under stronger-than-crash validation.
- Sep 25, 2026CorrectionLimiting what untrusted input can cause an agent to do gives injection resistance that does not depend on the model resisting.Willison's post quotes and builds on the design patterns paper, so it is not independent of it. Corroboration rests on separate organizations adopting the position, such as OpenAI's deterministic Lockdown Mode.
- Sep 25, 2026CorrectionChecking only that the original crash no longer reproduces overstates how often AI-generated patches actually fix the vulnerability.The review paper's support is its manual review of baseline agents (38-46% of fully validated patches semantically wrong), not the competition-scored accuracy figures.
- Sep 25, 2026CorrectionAgents under evaluation have coordinated through unintended shared channels, reused each other's artifacts, and tried to keep those channels alive.OpenAI's July 21 disclosure did not describe coordination. UK AISI (Aug 4) first reported agents reusing accounts and artefacts other agents left, and METR and OpenAI (Aug 26) described the message board; the cross-lab token reuse is OpenAI's account.
- Sep 25, 2026CorrectionCyber capability measured at fixed, low token budgets understates what frontier models can do and how fast they are improving.The OpenAI playbook cites UK AISI's own measurements, so all evidence comes from one evaluator.
- Sep 25, 2026CorrectionOn ExploitGym (May 2026), the strongest agents produced working exploits for 157 and 120 of 898 instances with mitigations off; with standard mitigations on, 45 and 21 survived.The pipeline audit covered eight knowledge and multiple-choice benchmarks, not ExploitGym. The qualification rests on ExploitBench, where no publicly deployed model reached code execution on V8.
- Sep 25, 2026CorrectionMalicious or compromised agent extensions, MCP servers, and skills have been published to public registries and used against real users.The Nx compromise was a malicious build-tool package that invoked installed AI CLIs, not a malicious agent extension, MCP server or skill. Corroboration rests on the malicious postmark-mcp server (found by Koi Security, disclosed by Postmark), independent of the Amazon Q incident.
- Sep 25, 2026CorrectionAgent monitors can be driven to miss covert actions: optimized attacks pushed suspicion scores near zero, and weak red-teaming overstates catch rates.SHADE-Arena and MonitoringBench share an author and count once. Corroboration rests on UK AISI's Control Red Team, which drove suspicion scores for covert API-key exposure to 3/100 on Google DeepMind and Anthropic monitors.
- Sep 25, 2026CorrectionScores on public CTF benchmarks can be inflated when agents find published solutions, and static benchmarks lose validity as their flaws are patched.The 2026-05-21 paper argues that benchmarks go stale but does not measure or independently test contamination, so the measured part of this claim rests on CTFusion alone.
- Sep 25, 2026CorrectionLLM agents fall well short of reliable performance on realistic threat-investigation and threat-hunting benchmarks built from security logs.CyberSOCEval tests multiple-choice question answering, not agents on investigation or hunting. Corroboration rests on Simbian's Cyber Defense Benchmark, where the best of five models flagged 3.8% of malicious events in raw logs.
- Sep 25, 2026CorrectionPublished prompt-injection defenses report attack success cut to near zero, or under 10%, against most of the fixed attacks their authors tested.Adaptive-attack results narrow this finding rather than dispute it, since its scope is limited to fixed attacks. The earlier entry's figure was wrong: spotlighting peaked at 82.4% under adaptive attack on Gemini, not above 90%.
- Sep 25, 2026Initial corpus
The initial corpus was assembled and verified against its sources: 219 records, 52 findings, 11 research openings, 31 methods, and 7 key questions. See methodology for how it was built.