Chronicle/Defense & research

AgentDojo: an extensible environment for prompt injection attacks and defenses on LLM agents

DefenseBenchmarkSignificance assistant-drafted

Debenedetti and colleagues (ETH Zurich, Invariant Labs) release AgentDojo, a dynamic environment with 97 realistic user tasks across workspace, banking, travel and Slack suites and 629 security test cases. It measures both utility and targeted attack success, and reports that existing attacks break some security properties but not all. It became the standard testbed used by CaMeL, US AISI/CAISI, LlamaFirewall and adaptive-attack studies.

Why it matters

Most later agent prompt-injection defense claims, and the adaptive attacks against them, are reported on AgentDojo.

Key facts

As stated in the sources, with where to find them.

  • 97 tasks and 629 security test cases.Abstract
  • GPT-4o: 69.00% benign utility, 50.08% utility under the 'important instructions' attack, 47.69% targeted attack success rate.Section 4.1, Figure 6; Appendix C Table 3 (current HTML version)
  • Targeted ASR for GPT-4o with defenses: tool filter 6.84%, prompt injection detector 7.95%, repeat user prompt 27.82%, data delimiting 41.65%, against 57.69% with no defense in the same table (current version).Section 4.3, Figure 9; Appendix C Table 5

Findings that cite this record

Key questions this bears on

Sources

Related records