Wallace and co-authors at OpenAI argue that models treat system prompts and untrusted inputs with equal priority and propose an explicit instruction hierarchy that tells the model which instructions to follow when they conflict. Applied to GPT-3.5, they report large robustness gains against attack types not seen in training with minimal capability loss.
Why it matters
The instruction hierarchy became OpenAI's stated foundation for prompt-injection robustness in later agent products.
Key facts
As stated in the sources, with where to find them.
- Training GPT-3.5 with the hierarchy is reported to drastically increase robustness, including to attack types not seen during training, with minimal degradation of standard capabilities.Abstract
Findings that cite this record
Sources
Related records
Mar 5, 2024
Mar 10, 2026
Oct 10, 2025
Feb 23, 2023
Jun 19, 2024
Nov 7, 2025