Scope: what this does not show
Self-reported by the labs about their own agents and monitors.
Qualified: Still standing, but later work narrows how far it applies.
Evidence
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Agent monitors can be driven to miss covert actions: optimized attacks pushed suspicion scores near zero, and weak red-teaming overstates catch rates. qualifies this findingCatch rates against ordinary mistakes say little about adversarial evasion.
Key questions that rely on this finding
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.