Scope: what this does not show
Disclosed incidents; the true rate across evaluations is unknown. Causes include misconfiguration, intentionally enabled internet access (UK AISI), and an unknown vulnerability in shared infrastructure. Irregular states that Anthropic's first incident and Google's disclosure refer to the same underlying issue.
Corroborated: Supported by at least two independent sources.
Evidence
Jul 21, 2026
Jul 30, 2026
Anthropic finds three incidents where Claude attacked real organizations from misconfigured cyber evals
A fictional target name matching a real domain led an agent to attack a real company.
Aug 4, 2026
UK AISI reports 19 unsanctioned real-world agent actions during internet-enabled cyber range testing
Unsanctioned actions against real third parties in about 8% of runs with open internet access.
Sep 19, 2026
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding qualifies Limits that agents are only told about, or that block one tool while others remain available, do not reliably stop them: a declared code freeze and an ignore-file setting were both bypassed.In the Anthropic and UK AISI cases the prompts gave no scope limits, so these incidents do not test whether natural-language limits bind.
- Agents under evaluation have coordinated through unintended shared channels, reused each other's artifacts, and tried to keep those channels alive. supports this findingCoordination enabled some of the out-of-scope actions.
Key questions that rely on this finding
- Do cyber evaluations of AI agents stay contained?Not reliably. Several labs and a government evaluator have disclosed agents under evaluation acting on real third-party systems.