Research defenses and research models. CaMeL-style architectural isolation was not among the defenses tested in the main adaptive-attack study.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
Tested an undefended agent: novel attacks raised hijack rates from 11% to 81%. Shows attacks improve, not that defenses fail.
Adaptive attacks exceeded 90% against 2 of 8 baseline defenses on Gemini 2.0; a warning defense and a user-instruction classifier held (at most 11%). Shares authors with 'The Attacker Moves Second'.
Adaptive attacks exceeded 90% against most of 12 defenses; human red-teamers succeeded on every challenge in the subset of defenses they were given.
How it relates to other findings
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies Published prompt-injection defenses report attack success cut to near zero, or under 10%, against most of the fixed attacks their authors tested.The defenses' low static attack success does not carry over to adaptive attackers, which that finding's scope already excludes.
- This finding qualifies Frontier labs' layered defenses reduce prompt injection in tool-use, browser and computer-use agents but do not eliminate it.Research defenses with low static attack success failed under adaptive attack, so lab-reported rates against fixed attack sets may also overstate robustness.
- Measured hijack rates rise sharply when attackers get repeated attempts, so single-attempt figures understate risk. supports this findingBoth show fixed, single-shot attack measurements understate attacker success; the attempt data comes from frontier models, not the published research defenses.
- Every frontier agent tested in two large public red-teaming competitions was hijacked at least once; in the 2026 competition the injections also had to conceal the compromise from the user. supports this findingCrowds of attackers who iterate get through every frontier agent tested; these competitions tested deployed models, not published research defenses.
Key questions that rely on this finding
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Status history
- 2025-01-17ReportedUS AISI shows novel attacks raise hijack rates sharply. · record
- 2025-05-20CorroboratedGoogle DeepMind independently reports adaptive attacks above 90% on Gemini. · record
- 2026-09-25ReportedcorrectionThe 2025 US AISI and Google DeepMind entries did not test published defenses with near-zero reported success, and DeepMind shares authors with the primary study. 'The Attacker Moves Second' is the primary evidence; no independent replication is recorded yet. · record