Findings/constrain-what-untrusted-input-can-trigger

Limiting what untrusted input can cause an agent to do gives injection resistance that does not depend on the model resisting.

Corroboratedargued5 evidence records from 4 independent sourcesassistant-drafted
Scope: what this does not show

A design position that several parties converge on. It trades away capability and is not a measurement.

Corroborated: Supported by at least two independent sources.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2025-06-10ReportedDesign patterns paper argues for constraining agents. · record
  2. 2025-06-16CorroboratedIndependent framing of the same principle as the lethal trifecta. · record
  3. 2026-09-25CorroboratedcorrectionWillison's post quotes and builds on the design patterns paper, so it is not independent of it. Corroboration rests on separate organizations adopting the position, such as OpenAI's deterministic Lockdown Mode. · record