Organizations/academic

Shandong University

1 records1 defenseWebsite
Sep 29, 2026
ToolFence proposes typed capabilities with parameter provenance to authorize agent tool calls, tested on AgentDojo
DefensePaperHong Kong Polytechnic University, Shandong University

Li, He, Dai and Xiao (arXiv v1, 29 September 2026) propose ToolFence, an inference-time defense that compiles the authenticated user request into typed capabilities with provenance constraints on authority-sensitive arguments, enforces them with a deterministic monitor, and asks an LLM judge to grant new capability shapes rather than judge each call. On AgentDojo with Qwen3-max the authors report overall attack success of 0.20% against 21.20% undefended, with clean utility 38.90% against 42.70% undefended; a cross-session cache of approved shapes cut judge calls per task from 1.84 to 1.05 in their ablation. The attacks are six injection strategies averaged; the authors report no adaptive attacks against ToolFence and residual failures inside authorized data flows.