NIST's AI safety institute technical staff (renamed the Center for AI Standards and Innovation in June 2025) extended AgentDojo and red-teamed agents built on the upgraded Claude 3.5 Sonnet. On held-out Workspace tasks, attack success rose from 11% for the strongest baseline attack to 81% for the strongest newly developed attack, and across five injection tasks from 57% to 80% when each attack was tried 25 times. The team released an Inspect-based AgentDojo port and ran the red teaming with the UK AI Security Institute.
Why it matters
A government evaluator showed that agent-hijacking scores depend heavily on attack novelty and attempt count, not only on the model.
Key facts
As stated in the sources, with where to find them.
- Claude 3.5 Sonnet (Oct 2024), held-out Workspace tasks: 11% ASR for the strongest AgentDojo baseline attack vs 81% for the strongest new red-team attack.Insight #2 section
- Across five injection tasks, average ASR was 57% at one attempt and 80% with 25 attempts per attack.Insight #3 and #4 sections
Findings that cite this record
Jan 17, 2025
Jan 17, 2025
Mar 5, 2024
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Mar 24, 2025
Feb 5, 2026
Mar 16, 2026
Oct 10, 2025
Jun 19, 2024
May 20, 2025