Desk/2024-W41

Week of Oct 7–13, 2024

1 records0 status changes on new evidence0 new findings

Defense & research

Oct 7, 2024
SecAlign uses preference optimization to train LLMs against prompt injection
DefensePaperMeta, UC Berkeley

Chen and colleagues (UC Berkeley and Meta) train models with preference optimization to prefer responses that follow the legitimate instruction over those that follow injected instructions. In the ACM CCS 2025 version they report injection success rates below 10% even for attacks more sophisticated than those seen in training, with utility similar to the undefended model; the October 2024 first version reported GCG-based injection success on Mistral-7B falling from 56% to 2%.