Chronicle/Defense & research

US AISI (later CAISI) shows red-team attacks and repeated attempts raise agent hijacking rates on AgentDojo

DefenseEvaluation reportSignificance assistant-drafted

NIST's AI safety institute technical staff (renamed the Center for AI Standards and Innovation in June 2025) extended AgentDojo and red-teamed agents built on the upgraded Claude 3.5 Sonnet. On held-out Workspace tasks, attack success rose from 11% for the strongest baseline attack to 81% for the strongest newly developed attack, and across five injection tasks from 57% to 80% when each attack was tried 25 times. The team released an Inspect-based AgentDojo port and ran the red teaming with the UK AI Security Institute.

Why it matters

A government evaluator showed that agent-hijacking scores depend heavily on attack novelty and attempt count, not only on the model.

Key facts

As stated in the sources, with where to find them.

  • Claude 3.5 Sonnet (Oct 2024), held-out Workspace tasks: 11% ASR for the strongest AgentDojo baseline attack vs 81% for the strongest new red-team attack.Insight #2 section
  • Across five injection tasks, average ASR was 57% at one attempt and 80% with 25 attempts per attack.Insight #3 and #4 sections

Findings that cite this record

Key questions this bears on

Sources

Related records