Nasr, Carlini, Tramèr and 11 co-authors apply gradient, reinforcement learning, search and human red-teaming attacks to 12 published defenses. Most defenses originally reported near-zero attack success, but the adaptive attacks exceed 90% success against most, and human red-teamers succeeded on every challenge in the subset of defenses they were given.
SecAlign
Preference-optimization defense against prompt injection, later released as Meta SecAlign models.
Records citing SecAlign
Oct 10, 2025
'The Attacker Moves Second': adaptive attacks bypass 12 published jailbreak and injection defenses
Oct 7, 2024
SecAlign uses preference optimization to train LLMs against prompt injection
Chen and colleagues (UC Berkeley and Meta) train models with preference optimization to prefer responses that follow the legitimate instruction over those that follow injected instructions. In the ACM CCS 2025 version they report injection success rates below 10% even for attacks more sophisticated than those seen in training, with utility similar to the undefended model; the October 2024 first version reported GCG-based injection success on Mistral-7B falling from 56% to 2%.