Google DeepMind introduces an AI Control Roadmap for internally deployed agents. It treats potentially misaligned AI as an insider threat, builds a threat model on MITRE ATT&CK tactics and techniques, and scales detection and response requirements with levels of agent evasion (D1-D4) and attack capability (R1-R3). It reports analyzing about a million coding agent tasks with a prototype asynchronous monitor, which informed a live monitor for its Gemini Spark agent. It says the majority of flagged events were not adversarial and often stemmed from misinterpretation or overeagerness.
Why it matters
A frontier lab is applying security-operations concepts directly to monitoring its own agents.
Key facts
As stated in the sources, with where to find them.
- Capability levels: detection evasion D1-D4 and attack execution R1-R3; metrics tracked include coverage, recall and time-to-response.Understanding AI Control; Scaling security as AI gets smarter
- About a million coding agent tasks analyzed by an internal asynchronous monitoring prototype; Google DeepMind says the majority of flagged events did not stem from adversarial intent and often reflected misinterpretation or overeagerness.Analyzing a million agent trajectories to inform live monitoring
Findings that cite this record
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Jul 23, 2026
Jun 17, 2025
Jun 9, 2026
May 15, 2026
May 10, 2026
Sep 2, 2026