Zou and colleagues (Gray Swan and collaborators; Anthropic describes the resulting benchmark as developed with the UK AI Security Institute) report a public red-teaming competition with 1.8 million prompt-injection attacks against 22 frontier agents in 44 deployment scenarios, producing over 60,000 successful policy violations. From these they build the Agent Red Teaming (ART) benchmark and find nearly all agents break within 10 to 100 queries for most behaviors, with high transfer and little correlation between robustness and model size or capability.
Why it matters
The ART benchmark it created is used by labs, including in Anthropic system cards, to report agent prompt-injection robustness.
Key facts
As stated in the sources, with where to find them.
- 1.8 million prompt-injection attacks, 22 agents, 44 scenarios, over 60,000 successful policy violations.Abstract
- ART benchmark evaluation of 19 models: most behaviors elicited within 10-100 queries; limited correlation of robustness with size, capability or inference-time compute.Abstract
Findings that cite this record
Jan 17, 2025
Jan 17, 2025
Jul 28, 2025
Jul 28, 2025
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Mar 16, 2026
Feb 5, 2026
Sep 30, 2025
Jun 10, 2025
Nov 24, 2025
Mar 24, 2025