Zhan, Liang, Ying and Kang release InjecAgent, a benchmark of 1,054 test cases spanning 17 user tools and 62 attacker tools, covering direct harm to users and exfiltration of private data. They evaluate 30 LLM agents and find a ReAct-prompted GPT-4 agent vulnerable in about a quarter of cases.
Why it matters
It was an early systematic measurement showing that tool-using agents follow instructions embedded in tool outputs.
Key facts
As stated in the sources, with where to find them.
- Benchmark has 1,054 test cases, 17 user tools and 62 attacker tools; 30 LLM agents evaluated.Abstract
- ReAct-prompted GPT-4 was vulnerable to the attacks 24% of the time; a reinforced 'hacking prompt' nearly doubled attack success.Abstract
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Jun 19, 2024
Nov 3, 2023
Aug 6, 2025
Jul 8, 2025
May 26, 2025
Feb 23, 2023