Scope: what this does not show
Depends on attack budget and model; figures come from lab and government red-teaming.
Corroborated: Supported by at least two independent sources.
Evidence
Jan 17, 2025
US AISI (later CAISI) shows red-team attacks and repeated attempts raise agent hijacking rates on AgentDojo
Claude 3.5 Sonnet, five new injection tasks: average success rose from 57% at one attempt to 80% with 25 attempts.
Feb 5, 2026
Claude Opus 4.6 system card reports prompt injection rates by surface, attempts and safeguards
Computer use, Opus 4.6 with extended thinking and no safeguards: 17.8% at 1 attempt versus 78.6% at 200 attempts.
Jul 28, 2025
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- This finding supports Attackers who adapt to a defense defeat most published prompt-injection defenses that reported near-zero success against static attacks.Both show fixed, single-shot attack measurements understate attacker success; the attempt data comes from frontier models, not the published research defenses.
- This finding qualifies Frontier labs' layered defenses reduce prompt injection in tool-use, browser and computer-use agents but do not eliminate it.Where labs quote single-attempt rates, they understate a persistent attacker: Opus 4.6 computer-use success rose from under 18% at 1 attempt to up to 78.6% at 200.