Huawei authors propose a plug-in layer in which fine-tuned 3B-parameter simulators, built on an uncensored model, draft a tree of the target agent's possible future actions, prune it against the actions the agent really takes, and pass the leaves to an existing detector. On AgentDojo with a Qwen3-235B agent they report attack success falling to 0% with two detectors, and on ActorAttack multi-turn jailbreaks to 0.0% at a sampling budget of 5. The attacks were benchmark attacks not designed against the defense, and no attacker aware of it was tested.
It reports a way to add forward-looking checks to existing detectors for indirect prompt injection and multi-turn jailbreaks, with benchmark results that have not been tested against an attacker aware of the defense.
Key facts
As stated in the sources, with where to find them.
- Setup: target agent Qwen3-235B-A22B-Instruct-2507; the user and assistant simulators are fine-tuned Dolphin3-3B models without safety alignment, deployed as Multi-LoRA adapters; baselines include Sandwich, Spotlight, Tool Filter, ProtectAI and PromptGuard classifiers, and a Qwen3-0.6B judge. The authors state the simulators add 1.3% parameter overhead relative to the 235B target.v1, Section 4.1; Limitations
- AgentDojo (four suites, six attack types), speculation depth 4 and width 8 (the Table 1 caption gives the width as M = 4 while the text says 8), average over all cells: no defense 31.78% attack success rate and 56.59% utility under attack; PromptGuard alone 12.25% and 36.02%, with SSH 0.00% and 44.14%; Qwen3-0.6B judge alone 10.48% and 44.75%, with SSH 0.00% and 58.43%. When a speculated branch scores as risky, the agent action is resampled up to five times and then generation is interrupted.v1, Section 4.2.1; Table 1; Table 6
- ActorAttack multi-turn jailbreaks, attacked by the uncensored Dolphin-X1-8B: risk exposure 84.5% at one sample on the HarmBench subset, and 0.0% attack success on both subsets at a sampling budget of 5. Speculation hit rate against the agent's real action was 47.2% on AgentDojo at budget 4 but 8% on ActorAttack.v1, Section 4.2.2; Table 2; Section 4.2.6
- On benign utility checks with the judge detector, BFCL-v3 overall accuracy was 45.13 without and 44.88 with SSH, XSTest refusal rate was 0.0% in both, and no benign sample on those two checks triggered an alert; separately, the system-level false-positive rate on AgentDojo, from the limitations section, was 0.2%, and a false positive adds roughly 3 to 7 seconds.v1, Section 4.2.3; Table 3; Limitations
- With GPT-4o-mini and GPT-OSS-120B as target agents on AgentDojo, average attack success was 8.4% and 14.9% without SSH and 0.0% and 0.7% with SSH plus the judge detector; the authors call this near-zero.v1, Appendix D.1; Table 5
- The 4,700-sample fine-tuning data for injection scenarios were synthesized from 30 attack templates; the paper does not report how closely these resemble the benchmarks' attacks, and it does not test attackers that know SSH is deployed.v1, Appendix B.3; Section 4
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.