Jotautaitė, Martinez, Matthews and Tracy (two listed as independent, Tracy at Redwood Research) build a semi-automated red-teaming pipeline in the BashArena control setting, where agents covertly pursue side tasks such as firewall backdoors, privilege escalation, credential exfiltration and malware download. MonitoringBench's 2,644 attack trajectories lowered a Claude Opus 4.5 monitor's catch rate at 1% false positives from 94.9% on directly elicited attacks to 60.3% on refined ones.
Why it matters
It suggests monitor evaluations that rely only on model-elicited attacks can overstate how well monitors catch security-relevant sabotage.
Key facts
As stated in the sources, with where to find them.
- 2,644 attack trajectories; Opus 4.5 monitor catch rate at a 1% false-positive threshold falls from 94.9% (elicited-only Opus attacks) to 60.3% (best refined attacks).Abstract
- 81% of Opus 4.6-elicited attacks fell into 2 of 5 technical tiers and 64% followed one structural template.Introduction; Section 3
- On the final benchmark, the Opus 4.5 monitor caught roughly 70% of attacks with full-trajectory monitoring and roughly 83% with stepwise monitoring; the authors say these are not directly comparable because stepwise thresholds give an effective 6.7-17.4% trajectory-level false-positive rate.Section 4, final MonitoringBench results (Figure 3)
Findings that cite this record
Key questions this bears on
- Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Jul 21, 2026
Mar 19, 2026
Jul 30, 2026
Oct 22, 2025
Jun 18, 2026
Jun 3, 2026