Findings/static-defense-results-hold

Published prompt-injection defenses report attack success cut to near zero, or under 10%, against most of the fixed attacks their authors tested.

Revalidatemeasured4 evidence records from 4 independent sourcesassistant-drafted
Scope: what this does not show

Results against fixed or known attacks, as reported by each defense's authors. StruQ's authors report 9% (TAP) and 58% (GCG) success on Llama for their strongest optimization attacks, and the instruction-hierarchy paper reports robustness gains rather than attack success rates. Says nothing about attackers who adapt to the defense.

Revalidate: Older than its half-life with no newer evidence. May no longer hold.

Evidence

How it relates to other findings

supportsqualifiescontestssupersedes
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic

Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.

Status history

  1. 2024-02-09ReportedStruQ reports reduced injection success. · record
  2. 2024-03-20CorroboratedMicrosoft reports spotlighting cuts injection success below 2% on static attacks. · record
  3. 2025-05-20ContestedGoogle DeepMind finds adaptive attacks exceed 90% success against spotlighting-style defenses on Gemini. · record
  4. 2026-09-25QualifiedcorrectionAdaptive-attack results narrow this finding rather than dispute it, since its scope is limited to fixed attacks. The earlier entry's figure was wrong: spotlighting peaked at 82.4% under adaptive attack on Gemini, not above 90%. · record
  5. 2026-09-26RevalidateComputed: 494 days since the last evidence, past the 365-day half-life.