Meta proposes that, within a session, an agent should have at most two of three properties: processing untrustworthy inputs, accessing sensitive systems or private data, and changing state or communicating externally. If all three are needed, the agent should not act autonomously and needs human approval or other validation. Meta illustrates this with travel, research and internal coding agent examples.
Why it matters
It turns the lethal trifecta idea into an explicit design rule that a major platform company endorses.
Key facts
As stated in the sources, with where to find them.
- Properties: [A] untrustworthy inputs, [B] sensitive systems or private data, [C] changing state or external communication; satisfy no more than two per session.Framework section
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Oct 8, 2025
Sep 25, 2025
Jan 28, 2026
Feb 13, 2026
Jan 19, 2026
Aug 20, 2025