OpenAI describes IH-Challenge, a reinforcement learning dataset of simple, programmatically graded conflicts between higher- and lower-privilege instructions designed to avoid shortcuts such as over-refusal. A GPT-5 Mini variant trained on it (GPT-5 Mini-R) improved on instruction-hierarchy benchmarks and on CyberSecEval 2 and an internal prompt injection benchmark, with little capability loss; the dataset is publicly released.
Why it matters
It is an open training resource for model-level prompt injection robustness from a frontier lab.
Key facts
As stated in the sources, with where to find them.
- GPT-5 Mini vs GPT-5 Mini-R: TensorTrust (dev-user) 0.76 to 0.91; Developer<>User conflict 0.83 to 0.95; IH-Challenge over-refusal 0.79 to 1.00; GPQA Diamond unchanged at 0.83.Results tables
- Prompt injection robustness improved on CyberSecEval 2 and an internal static benchmark; exact scores shown only in charts.Prompt injection robustness section
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Oct 10, 2025
Sep 30, 2025
Aug 4, 2026
Dec 22, 2025
Apr 19, 2024
Feb 13, 2026