Nasr, Carlini, Tramèr and 11 co-authors apply gradient, reinforcement learning, search and human red-teaming attacks to 12 published defenses. Most defenses originally reported near-zero attack success, but the adaptive attacks exceed 90% success against most, and human red-teamers succeeded on every challenge in the subset of defenses they were given.
Why it matters
It is the central evidence that static-benchmark robustness claims for prompt injection defenses do not hold against adaptive attackers.
Key facts
As stated in the sources, with where to find them.
- 12 defenses bypassed with attack success above 90% for most; most had originally reported near-zero ASR.Abstract
- Spotlighting and Prompt Sandwiching on AgentDojo: as low as 1% attack success under the benchmark's static attacks (authors' re-implementation) vs over 95% with the adaptive search attack.Section 5.1
- Meta SecAlign (AgentDojo): 2% originally vs 96% adaptive; PromptGuard and Protect AI detector: over 90%; PIGuard: 71%; MELON: 76%, rising to 95% with full defense knowledge.Sections 5.2-5.4
- A human red-teaming competition with 500+ participants and a $20,000 prize pool succeeded in every challenge on the subset of defenses included in the human study.Section 6; Appendix E
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
May 20, 2025
May 6, 2025
Jan 17, 2025
Mar 10, 2026
Feb 5, 2026
Nov 24, 2025