The UK AI Security Institute, with Redwood Research, releases ControlArena, an open-source library built on Inspect for running AI control experiments. It bundles settings from simple programming problems to infrastructure-as-code codebases, attack policies, monitors and protocols such as trusted editing and defer-to-trusted, and AISI says researchers at Anthropic, Google DeepMind and Redwood have used it.
Why it matters
It standardizes testbeds for measuring whether monitors and protocols stop an agent pursuing covert harmful side tasks.
Key facts
As stated in the sources, with where to find them.
- Includes monitors that score suspicion, trusted editing, defer-to-trusted micro-protocols and macro-protocols across tasks.Blog, protocols section
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Jul 21, 2026
Mar 1, 2026
May 10, 2026
Aug 18, 2026
Jun 17, 2025
Jul 23, 2026