Greenblatt, Shlegeris, Sachan and Roger propose evaluating safety protocols against an untrusted model that is deliberately trying to subvert them. In a programming testbed, GPT-4 acts as the untrusted model, GPT-3.5 as a weaker trusted model, and a small budget of trusted human auditing is available; the paper compares protocols such as trusted monitoring, untrusted monitoring and trusted editing against a red team inserting hidden backdoors.
Why it matters
It founded the AI control framing that later monitoring work on agents (ControlArena, UK AISI's Control Red Team, lab coding-agent monitors) builds on.
Key facts
As stated in the sources, with where to find them.
- With a 2% auditing budget, trusted monitoring scored 62% safety at 97% usefulness; upfront auditing alone scored 15% safety at 100% usefulness.Table 2 / Figure 2
- Trusted editing (20% threshold) scored 92% safety at 94% usefulness; untrusted monitoring scored 79% safety at 96% usefulness.Table 2; Section 3.5
Findings that cite this record
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Jun 17, 2025
May 6, 2025
May 10, 2026
Oct 22, 2025
Apr 19, 2024
Jun 18, 2026