NIST released the 2025 edition of its adversarial machine learning taxonomy, co-authored with the UK AI Security Institute and US AI Safety Institute staff. Unlike the 2023 edition, it includes a section on the security of agents, noting that tool-using agents are exposed to direct and indirect prompt injection and that hijacking can lead to arbitrary code execution or data exfiltration.
Why it matters
It is the reference US government taxonomy that COSAiS overlays and CAISI agent work build on.
Key facts
As stated in the sources, with where to find them.
- Section 3.5 'Security of Agents' states agents are vulnerable to direct and indirect prompt injection and that tool use lets attackers hijack agents to execute arbitrary code or exfiltrate data.Section 3.5
- Section 3.6 cites AgentDojo as a framework for measuring agent vulnerability to prompt injection via tool-returned data, and AgentHarm among jailbreak benchmarks.Section 3.6 Benchmarks
- The E2023 edition had no dedicated agents section (its prompt injection coverage was Sections 3.3-3.4).E2023 table of contents
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Jan 17, 2025
Mar 24, 2025
Jan 8, 2026
May 22, 2025
Mar 16, 2026
Aug 20, 2025