Findings/lab-defenses-reduce-not-eliminate

Frontier labs' layered defenses reduce prompt injection in tool-use, browser and computer-use agents but do not eliminate it.

Qualifiedreported5 evidence records from 3 independent sourcesassistant-drafted
Scope: what this does not show

Rates are self-reported by the labs that built the systems; no independent adaptive test of 2026 production defenses is recorded.

Qualified: Still standing, but later work narrows how far it applies.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Key questions that rely on this finding

Status history

  1. 2025-05-20ReportedGoogle DeepMind: adversarial fine-tuning reduced but did not eliminate success. · record
  2. 2025-08-25CorroboratedAnthropic reports residual success after mitigations in its browser agent. · record
  3. 2026-02-05QualifiedMulti-attempt figures in Anthropic's system card are far higher than single-attempt rates. · record