2024-03-20
Published prompt-injection defenses report attack success cut to near zero, or under 10%, against most of the fixed attacks their authors tested.
Microsoft reports spotlighting cuts injection success below 2% on static attacks.
Microsoft reports spotlighting cuts injection success below 2% on static attacks.
Hines and colleagues at Microsoft describe spotlighting, a family of prompt-engineering transformations (delimiting, datamarking, encoding) that signal to the model where untrusted text came from. On GPT-family models they report attack success falling from above 50% to under 2% with minimal task impact. Microsoft later described spotlighting as one layer of its production defense-in-depth.