Findings/production-monitors-catch-escalated-incidents

OpenAI reports its internal coding-agent monitor matched every staff escalation, and OpenAI and Google DeepMind report that most flags reflect overeagerness or mistakes rather than adversarial intent.

Qualifiedreported2 evidence records from 2 independent sourcesassistant-drafted
Scope: what this does not show

Self-reported by the labs about their own agents and monitors.

Qualified: Still standing, but later work narrows how far it applies.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2026-03-19ReportedOpenAI reports its monitor matched every staff-escalated incident. · record
  2. 2026-07-23QualifiedUK AISI shows optimized attacks can drive monitor suspicion near zero. · record