Scope: what this does not show
Red-team settings designed to stress monitors.
Corroborated: Supported by at least two independent sources.
Evidence
Jun 17, 2025
May 10, 2026
MonitoringBench shows refined covert attacks cut an Opus 4.5 monitor's catch rate from 95% to 60%
Model-elicited attacks are mode-collapsed and overstate catch rates.
Jul 23, 2026
UK AISI Control Red Team finds evasion weaknesses in Google DeepMind and Anthropic agent monitors
Suspicion scores for covert API-key exposure driven down to 3 out of 100.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies OpenAI reports its internal coding-agent monitor matched every staff escalation, and OpenAI and Google DeepMind report that most flags reflect overeagerness or mistakes rather than adversarial intent.Catch rates against ordinary mistakes say little about adversarial evasion.
Key questions that rely on this finding
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Status history
- 2025-06-17ReportedSHADE-Arena: some hidden sabotage goes undetected. · record
- 2026-05-10CorroboratedMonitoringBench shows weak attacks overstate monitor performance. · record
- 2026-09-25CorroboratedcorrectionSHADE-Arena and MonitoringBench share an author and count once. Corroboration rests on UK AISI's Control Red Team, which drove suspicion scores for covert API-key exposure to 3/100 on Google DeepMind and Anthropic monitors. · record