Chronicle/Defense & research

OpenAI trains models to prioritize privileged instructions via an instruction hierarchy

DefensePaperSignificance assistant-drafted

Wallace and co-authors at OpenAI argue that models treat system prompts and untrusted inputs with equal priority and propose an explicit instruction hierarchy that tells the model which instructions to follow when they conflict. Applied to GPT-3.5, they report large robustness gains against attack types not seen in training with minimal capability loss.

Why it matters

The instruction hierarchy became OpenAI's stated foundation for prompt-injection robustness in later agent products.

Key facts

As stated in the sources, with where to find them.

  • Training GPT-3.5 with the hierarchy is reported to drastically increase robustness, including to attack types not seen during training, with minimal degradation of standard capabilities.Abstract

Findings that cite this record

Sources

Related records