Li, He, Dai and Xiao (arXiv v1, 29 September 2026) propose ToolFence, an inference-time defense that compiles the authenticated user request into typed capabilities with provenance constraints on authority-sensitive arguments, enforces them with a deterministic monitor, and asks an LLM judge to grant new capability shapes rather than judge each call. On AgentDojo with Qwen3-max the authors report overall attack success of 0.20% against 21.20% undefended, with clean utility 38.90% against 42.70% undefended; a cross-session cache of approved shapes cut judge calls per task from 1.84 to 1.05 in their ablation. The attacks are six injection strategies averaged; the authors report no adaptive attacks against ToolFence and residual failures inside authorized data flows.
Hong Kong Polytechnic University
Sep 29, 2026
ToolFence proposes typed capabilities with parameter provenance to authorize agent tool calls, tested on AgentDojo
Sep 26, 2026
CyberClear benchmarks LLM agents on reconstructing APT attack chains from long defender logs
Chen and colleagues (arXiv v1, under review for ICLR 2027) introduce CyberClear, 450 instances built from public APT log datasets in which an agent must turn long defender logs, without prior attack clues, into a provenance graph with ATT&CK-mapped steps. References were generated by an LLM and passed automatic verification and expert review; scoring compares graph code with five LLM-judge dimensions. They also propose CyberProvenance, a multi-agent harness that reproduces predicted attack steps in isolated lab environments and refines the graph, and report it best on most semantic metrics while the highest Strict score remains 0.6477 out of 1.