Benchmark settings with the models and attacks of 2024 and early 2025; rates vary widely by model, attack, and task, and are much lower for frontier models tested in 2026.
Qualified: Still standing, but later work narrows how far it applies.
Evidence
ReAct-prompted GPT-4 followed injected instructions in 24% of base-setting cases.
Undefended GPT-4o executed the attacker's goal in about 48% of cases under one attack.
Indirect injection averaged 27.55% success across 13 backbones; a mixed attack adding direct injection and memory poisoning averaged 84.30%.
Red-team attacks on Claude 3.5 Sonnet (October 2024) agents reached 81% success on held-out AgentDojo tasks.
How it relates to other findings
Select a finding to see how it relates to others. Arrows point from the newer finding to the one it supports, qualifies, contests, or supersedes.
- Every frontier agent tested in two large public red-teaming competitions was hijacked at least once; in the 2026 competition the injections also had to conceal the compromise from the user. qualifies this findingOn 2025 and 2026 frontier models with their built-in safeguards, per-model success in the 2026 competition was 0.5% to 8.5%, far below 2024 benchmark rates, though no model was immune.
Key questions that rely on this finding
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Status history
- 2024-03-05ReportedInjecAgent measures agents following injected instructions. · record
- 2024-06-19CorroboratedAgentDojo, from a different group, measures the same failure. · record
- 2026-09-26QualifiedPublic red-teaming competitions on 2025 and 2026 frontier models with built-in safeguards report much lower per-model success (0.5% to 8.5% in 2026), though every model was hijacked at least once. The substantial rates describe 2024 models and benchmarks. · record