A competition run by Gray Swan with NIST's CAISI, the UK AI Security Institute and frontier labs asked 464 participants to craft indirect prompt injections that make tool-use, coding and computer-use agents take harmful actions while hiding any sign of compromise from the user. Participants made 272,000 attempts against 13 frontier models, yielding 8,648 successes; per-model success ranged from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro), and at least one attack succeeded against every model.
Why it matters
It adds concealment to the success criterion and finds that each of the 13 frontier models tested fell to at least one indirect injection.
Key facts
As stated in the sources, with where to find them.
- 464 participants, 272,000 attack attempts, 13 frontier models, 8,648 successful attacks.Abstract
- Per-model success ranged from 0.5% (Claude Opus 4.5) to 8.5% (Gemini 2.5 Pro); universal strategies transferred across 21 of 41 behaviors.Abstract
- CAISI reports at least one successful attack against every target model, and transfer tended to flow from more robust to less robust models.NIST blog, key findings
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Feb 5, 2026
Sep 30, 2025
Jul 28, 2025
May 1, 2026
Nov 24, 2025
Apr 21, 2026