Scope: what this does not show
Results against fixed or known attacks, as reported by each defense's authors. StruQ's authors report 9% (TAP) and 58% (GCG) success on Llama for their strongest optimization attacks, and the instruction-hierarchy paper reports robustness gains rather than attack success rates. Says nothing about attackers who adapt to the defense.
Revalidate: Older than its half-life with no newer evidence. May no longer hold.
Evidence
Feb 9, 2024
Mar 20, 2024
Apr 19, 2024
Oct 7, 2024
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Attackers who adapt to a defense defeat most published prompt-injection defenses that reported near-zero success against static attacks. qualifies this findingThe defenses' low static attack success does not carry over to adaptive attackers, which that finding's scope already excludes.
Status history
- 2024-02-09ReportedStruQ reports reduced injection success. · record
- 2024-03-20CorroboratedMicrosoft reports spotlighting cuts injection success below 2% on static attacks. · record
- 2025-05-20ContestedGoogle DeepMind finds adaptive attacks exceed 90% success against spotlighting-style defenses on Gemini. · record
- 2026-09-25QualifiedcorrectionAdaptive-attack results narrow this finding rather than dispute it, since its scope is limited to fixed attacks. The earlier entry's figure was wrong: spotlighting peaked at 82.4% under adaptive attack on Gemini, not above 90%. · record
- 2026-09-26RevalidateComputed: 494 days since the last evidence, past the 365-day half-life.