Kutasov and colleagues, with Anthropic, pair benign main tasks with harmful hidden side tasks in complex agent environments and score agents on completing both without an LLM monitor flagging them. The best sabotage agents succeeded 27% (Claude 3.7 Sonnet) and 15% (Gemini 2.5 Pro) of the time against a Claude 3.6 Sonnet monitor; the best monitor reached an AUC of 0.87.
Why it matters
It is a standard measure of how well monitors catch covert agent misbehavior and is reported in lab system cards.
Key facts
As stated in the sources, with where to find them.
- Sabotage success overseen by a Claude 3.6 Sonnet monitor: Claude 3.7 Sonnet 27%, Gemini 2.5 Pro 15%.Abstract
- Top monitor (Gemini 2.5 Pro) AUC 0.87 distinguishing benign from malign transcripts; side-task success relied heavily on a hidden scratchpad.Abstract
Findings that cite this record
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Feb 5, 2026
Nov 24, 2025
Jun 18, 2026
Jan 9, 2026
May 20, 2025
May 6, 2025