Chronicle/Defense & research

'The Attacker Moves Second': adaptive attacks bypass 12 published jailbreak and injection defenses

DefensePaperSignificance assistant-drafted

Nasr, Carlini, Tramèr and 11 co-authors apply gradient, reinforcement learning, search and human red-teaming attacks to 12 published defenses. Most defenses originally reported near-zero attack success, but the adaptive attacks exceed 90% success against most, and human red-teamers succeeded on every challenge in the subset of defenses they were given.

Why it matters

It is the central evidence that static-benchmark robustness claims for prompt injection defenses do not hold against adaptive attackers.

Key facts

As stated in the sources, with where to find them.

  • 12 defenses bypassed with attack success above 90% for most; most had originally reported near-zero ASR.Abstract
  • Spotlighting and Prompt Sandwiching on AgentDojo: as low as 1% attack success under the benchmark's static attacks (authors' re-implementation) vs over 95% with the adaptive search attack.Section 5.1
  • Meta SecAlign (AgentDojo): 2% originally vs 96% adaptive; PromptGuard and Protect AI detector: over 90%; PIGuard: 71%; MELON: 76%, rising to 95% with full defense knowledge.Sections 5.2-5.4
  • A human red-teaming competition with 500+ participants and a $20,000 prize pool succeeded in every challenge on the subset of defenses included in the human study.Section 6; Appendix E

Findings that cite this record

Key questions this bears on

Sources

Related records