Signals
Approval prompts have hidden the actual command, and one-time approvals could be changed afterward.
One deployment study found analysts reused an agent's drafts heavily (median 96.7% reuse across 108 coded tickets from six analysts).
Central to FID-076's comparison of approval rules and evidence-sensitive policies.
Why it matters
Human approval is the default control for consequential agent actions. If it is a rubber stamp, the control is weaker than it looks.
Hypothesis
Approval accuracy falls with approval volume and rises when the interface shows the consequential effect rather than the command.
A first study
A reviewer study with security practitioners approving simulated agent actions under different interfaces and loads.
Controls it would need
Consent and ethics review; synthetic actions; balanced benign and harmful cases.
What it could and could not claim
Would describe the tested interfaces and participants, not all approval workflows.