Scope: what this does not show
Rates are self-reported by the labs that built the systems; no independent adaptive test of 2026 production defenses is recorded.
Qualified: Still standing, but later work narrows how far it applies.
Evidence
May 20, 2025
Google DeepMind reports lessons from continuously attacking Gemini with adaptive prompt injections
Gemini 2.5 tool-use scenarios (email, calendar); adversarial training cut TAP success from 99.8% to 53.6% in the email scenario.
Aug 25, 2025
Anthropic publishes prompt injection red-team rates for its Claude in Chrome browser agent pilot
11.2% of internal red-team cases after mitigations.
Nov 24, 2025
Anthropic reports 1.4% prompt injection success for Claude Opus 4.5 with improved Chrome extension safeguards
About 1% adaptive-attacker success reported.
Dec 22, 2025
Feb 5, 2026
Claude Opus 4.6 system card reports prompt injection rates by surface, attempts and safeguards
Claude Opus 4.6 with extended thinking, stronger Shade attacker: safeguards cut 200-attempt computer-use success from 78.6% to 57.1%.
How it relates to other findings
ReportedCorroboratedQualifiedContestedSupersededRevalidate· node size = evidence records · columns group by topic
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Attackers who adapt to a defense defeat most published prompt-injection defenses that reported near-zero success against static attacks. qualifies this findingResearch defenses with low static attack success failed under adaptive attack, so lab-reported rates against fixed attack sets may also overstate robustness.
- Measured hijack rates rise sharply when attackers get repeated attempts, so single-attempt figures understate risk. qualifies this findingWhere labs quote single-attempt rates, they understate a persistent attacker: Opus 4.6 computer-use success rose from under 18% at 1 attempt to up to 78.6% at 200.
Key questions that rely on this finding
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Status history
- 2025-05-20ReportedGoogle DeepMind: adversarial fine-tuning reduced but did not eliminate success. · record
- 2025-08-25CorroboratedAnthropic reports residual success after mitigations in its browser agent. · record
- 2026-02-05QualifiedMulti-attempt figures in Anthropic's system card are far higher than single-attempt rates. · record