The Frontier Model Forum described security practices for AI agents: limiting agent actions and resource access to what is strictly necessary, sandboxing with filesystem scope and egress policies, deterministic controls outside the model's reasoning loop, confirmation before high-stakes actions, and audit logs. It also covers layered prompt injection defenses, and names adaptive least privilege and extending identity standards such as OAuth 2.0 to agents as promising or developing areas.
Why it matters
It documents what frontier developers say they actually do to contain their own agents.
Key facts
As stated in the sources, with where to find them.
- Describes deterministic controls outside the agent's reasoning loop, such as limits on tool access and permitted actions that cannot be overridden through context manipulation, as valuable; sandboxing, filesystem scope and network egress policy determine what an agent can reach.Issue brief body
- Describes requiring confirmation before high-stakes or irreversible actions and audit logs that support incident investigation; says oversight may shift to risk-tiered escalation in which humans intervene only for consequential, sensitive or low-confidence actions.Issue brief body
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Jan 8, 2026
May 15, 2026
May 1, 2026
Oct 20, 2025
Jan 31, 2025
Aug 4, 2026