Scope: what this does not show
Competition settings with many attackers; per-attempt success is low and varies about 17-fold across models.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
Jul 28, 2025
Mar 16, 2026
CAISI, UK AISI and Gray Swan competition finds concealed indirect injections succeed on all 13 frontier models
All 13 frontier models were hijacked at least once; per-attempt success ranged from 0.5% to 8.5%.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding supports Attackers who adapt to a defense defeat most published prompt-injection defenses that reported near-zero success against static attacks.Crowds of attackers who iterate get through every frontier agent tested; these competitions tested deployed models, not published research defenses.
- This finding qualifies Undefended tool-using agents follow injected instructions in a substantial share of benchmark cases.On 2025 and 2026 frontier models with their built-in safeguards, per-model success in the 2026 competition was 0.5% to 8.5%, far below 2024 benchmark rates, though no model was immune.
Key questions that rely on this finding
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.