{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/toolfence-capability-authorization-2026/",
 "asOf": "2026-10-01",
 "id": "toolfence-capability-authorization-2026",
 "date": "2026-09-29",
 "datePrecision": "day",
 "title": "ToolFence proposes typed capabilities with parameter provenance to authorize agent tool calls, tested on AgentDojo",
 "lane": "defense",
 "kind": "paper",
 "summary": "Li, He, Dai and Xiao (arXiv v1, 29 September 2026) propose ToolFence, an inference-time defense that compiles the authenticated user request into typed capabilities with provenance constraints on authority-sensitive arguments, enforces them with a deterministic monitor, and asks an LLM judge to grant new capability shapes rather than judge each call. On AgentDojo with Qwen3-max the authors report overall attack success of 0.20% against 21.20% undefended, with clean utility 38.90% against 42.70% undefended; a cross-session cache of approved shapes cut judge calls per task from 1.84 to 1.05 in their ablation. The attacks are six injection strategies averaged; the authors report no adaptive attacks against ToolFence and residual failures inside authorized data flows.",
 "whyItMatters": "It targets within-tool hijacking, where an injected instruction keeps the permitted tool but changes an argument such as the recipient, which tool-level allowlists cannot see, and reports what that authorization design costs in utility, judge calls and runtime.",
 "actors": [
  "hong-kong-polytechnic-university",
  "shandong-university"
 ],
 "topics": [
  "prompt-injection",
  "access-controls",
  "tool-and-mcp-security"
 ],
 "atlas": [
  "tools",
  "untrusted-content",
  "monitor"
 ],
 "artifacts": [
  "agentdojo",
  "camel",
  "spotlighting",
  "gpt-4-family"
 ],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.37196",
   "publisher": "arXiv",
   "title": "ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents",
   "date": "2026-09-29",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "Version: figures are from arXiv v1 (29 September 2026). Authors are affiliated with Hong Kong Polytechnic University and Shandong University.",
   "locator": "Title page"
  },
  {
   "fact": "Typed blueprint: before the agent reads any external content, an isolated LLM 'policy architect' that sees only the authenticated query and controller-owned tool schemas assigns each tool an effect label and per-parameter authority-sensitivity flags, and lists capabilities. A capability is a tool, its effect, parameter bindings, a call budget and a reusability flag. Bindings declare where an argument must come from: copied from the user request, copied from the output of a named source tool, matching an authenticated template, or free (not authority-sensitive). The blueprint records provenance constraints, not concrete values.",
   "locator": "Sec. 3.2"
  },
  {
   "fact": "Deterministic monitor: every proposed call is checked without an LLM for capability match, binding compliance and remaining budget, and dispatched immediately if all hold. Read-only calls whose arguments all come from the user request are auto-allowed; reads whose arguments derive from tool output are sent to the judge, because unconditionally trusting reads measured 60% attack success in the authors' pilot on AgentDojo's Slack injection task, which is completed by fetching an attacker-controlled URL (the pilot's model and size are not stated).",
   "locator": "Sec. 3.3"
  },
  {
   "fact": "Runtime capability grant: when a call misses the blueprint, the controller labels each argument as user, tool or model provenance and refuses the proposal if an authority-sensitive argument is not traceable to the user or a declared source tool. Otherwise the judge sees the capability shape but not the concrete values and grants or denies it; a grant extends the running blueprint, and a denial falls back to one per-call judge verdict.",
   "locator": "Sec. 3.4"
  },
  {
   "fact": "Cross-session cache: approved shapes are stored under a value-free signature (tool, effect, binding kinds and source-tool names), never argument values or evidence. On reuse in a new session every concrete argument is still checked against the declared source as executed in that session. The authors state a fail-closed property: a call outside the blueprint and not granted is not dispatched, and unparsable judge output is treated as denial.",
   "locator": "Sec. 3.5; Appendix A.1, A.2"
  },
  {
   "fact": "Evaluation design: AgentDojo's 97 user tasks and 27 injection tasks give 629 pairs, which the authors split by executable tool calls into 524 cross-tool, 85 within-tool (same tool, changed authority-sensitive argument) and 20 ambiguous pairs (mainly non-tool denial-of-service in Travel). Overall attack success is computed over the 609 non-ambiguous pairs, averaged over six fixed injection strategies. Models are Qwen3-max, GPT-4o (gpt-4o-2024-08-06) and, in the appendix, Mistral-Small-3.1-24B; baselines are Repeat Prompt, Spotlighting, Tool Filter, PromptArmor, SecInfer and CaMeL.",
   "locator": "Sec. 4.1; Appendix C.1"
  },
  {
   "fact": "Qwen3-max, Table 1 (overall / cross-tool / within-tool attack success; clean utility; utility under attack): no defense 21.20 / 20.08 / 28.11, 42.70, 31.85; ToolFence 0.20 / 0.10 / 0.80, 38.90, 32.45; CaMeL 0.42 / 0.20 / 1.80, 32.67, 25.82; Tool Filter 1.67 / 0.45 / 9.18, 30.44, 31.88; PromptArmor 7.09 overall; Repeat Prompt 10.95; SecInfer 14.31; Spotlighting 18.56. Values are point estimates in percent; the table gives no confidence intervals.",
   "locator": "Sec. 4.2; Table 1"
  },
  {
   "fact": "GPT-4o, Table 1: no defense 38.01 overall, clean utility 84.20, utility under attack 51.00; ToolFence 0.90 overall (0.70 cross-tool, 2.10 within-tool), 82.70, 73.80; CaMeL 1.28 overall, 70.50, 47.80; Tool Filter 6.76 overall (4.90 cross-tool, 18.20 within-tool). Mistral-Small-3.1-24B, Table 4: no defense 13.98 overall, clean utility 50.39; ToolFence 0.13 (0.05 cross-tool, 0.60 within-tool), clean utility 55.80, utility under attack 52.30.",
   "locator": "Table 1; Table 4 (Appendix C.3)"
  },
  {
   "fact": "Ablation on Qwen3-max, Table 2 (utility under attack; attack success; judge calls per task; runtime multiple): per-call runtime judge 31.20, 4.35%, 6.35, 3.20x; blueprint plus deterministic monitor 25.60, 0.45%, 0.00, 1.45x; plus runtime capability grant 30.40, 0.18%, 1.84, 2.78x; plus cross-session cache 31.80, 0.18%, 1.05, 1.96x; full ToolFence with read auto-allow 32.45, 0.20%, 0.82, 1.90x.",
   "locator": "Sec. 4.3; Table 2"
  },
  {
   "fact": "The '43% fewer judge calls' figure is the cross-session cache step of the Qwen3-max ablation: judge invocations per task fall from 1.84 to 1.05 (a 42.9% reduction) with security and utility essentially unchanged. The abstract-level and Sec. 3.5 statements give the figure without that scope; the paper says it measured cache-hit rate but the archived text reports no value for it.",
   "locator": "Sec. 4.3 (Effect of caching); Sec. 3.5; Sec. 4.1"
  },
  {
   "fact": "Runtime, mean end-to-end per sample over 50 clean tasks and 200 attacked pairs: ToolFence 17.1 s (3.79x) on GPT-4o and 10.3 s (1.63x) on Qwen3-max; CaMeL 48.0 s (10.67x) and 81.7 s (12.95x); SecInfer 4.51x and 3.77x. The text does not label which model each CaMeL and SecInfer figure belongs to (the arithmetic matches GPT-4o then Qwen3-max). The Qwen3-max ablation reports 1.90x for full ToolFence, and the paper does not reconcile it with 1.63x.",
   "locator": "Sec. 4.5; Table 2"
  },
  {
   "fact": "Stated limits: a static blueprint alone gave only 25.6% utility under attack; legitimate values available only through untrusted content (for example a URL inside a file) are not auto-allowed, which costs utility, and the authors suggest selective human confirmation as a mitigation without testing it. In one Qwen3-max Travel case all four calls were permitted because restaurant names came from an authorized list, yet an injection influenced which candidate the model chose; the authors say provenance-only enforcement checks where a value comes from, not whether the choice faithfully reflects user intent, and that non-zero failures remain.",
   "locator": "Sec. 4.2; Sec. 4.4"
  },
  {
   "fact": "Scope of evidence: simulated AgentDojo environments with six fixed injection strategies; the archived text reports no adaptive attacks against ToolFence and does not name the judge or policy-architect model. The authors say they will release evaluation artifacts with safeguards for dual-use research.",
   "locator": "Sec. 4.1; Ethics statement"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-074"
 ],
 "methods": [
  "ai-monitoring",
  "capability-restriction",
  "control-data-isolation",
  "indirect-prompt-injection",
  "injection-classifiers",
  "input-delimiting"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}