{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/speculative-safety-honeypot-multi-turn-agent-defense-2026/",
 "asOf": "2026-10-01",
 "id": "speculative-safety-honeypot-multi-turn-agent-defense-2026",
 "date": "2026-09-30",
 "datePrecision": "day",
 "title": "Speculative Safety Honeypot: Huawei authors predict an agent's next actions with small simulators to flag multi-turn attacks",
 "lane": "defense",
 "kind": "paper",
 "summary": "Huawei authors propose a plug-in layer in which fine-tuned 3B-parameter simulators, built on an uncensored model, draft a tree of the target agent's possible future actions, prune it against the actions the agent really takes, and pass the leaves to an existing detector. On AgentDojo with a Qwen3-235B agent they report attack success falling to 0% with two detectors, and on ActorAttack multi-turn jailbreaks to 0.0% at a sampling budget of 5. The attacks were benchmark attacks not designed against the defense, and no attacker aware of it was tested.",
 "whyItMatters": "It reports a way to add forward-looking checks to existing detectors for indirect prompt injection and multi-turn jailbreaks, with benchmark results that have not been tested against an attacker aware of the defense.",
 "actors": [
  "huawei"
 ],
 "topics": [
  "prompt-injection",
  "jailbreaks-and-safeguards",
  "monitoring-and-control"
 ],
 "atlas": [
  "monitor",
  "untrusted-content",
  "tools"
 ],
 "artifacts": [
  "agentdojo",
  "gpt-4-family"
 ],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.39549",
   "publisher": "arXiv",
   "title": "Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks",
   "date": "2026-09-30",
   "type": "primary",
   "accessed": "2026-10-01"
  }
 ],
 "keyFacts": [
  {
   "fact": "Setup: target agent Qwen3-235B-A22B-Instruct-2507; the user and assistant simulators are fine-tuned Dolphin3-3B models without safety alignment, deployed as Multi-LoRA adapters; baselines include Sandwich, Spotlight, Tool Filter, ProtectAI and PromptGuard classifiers, and a Qwen3-0.6B judge. The authors state the simulators add 1.3% parameter overhead relative to the 235B target.",
   "locator": "v1, Section 4.1; Limitations"
  },
  {
   "fact": "AgentDojo (four suites, six attack types), speculation depth 4 and width 8 (the Table 1 caption gives the width as M = 4 while the text says 8), average over all cells: no defense 31.78% attack success rate and 56.59% utility under attack; PromptGuard alone 12.25% and 36.02%, with SSH 0.00% and 44.14%; Qwen3-0.6B judge alone 10.48% and 44.75%, with SSH 0.00% and 58.43%. When a speculated branch scores as risky, the agent action is resampled up to five times and then generation is interrupted.",
   "locator": "v1, Section 4.2.1; Table 1; Table 6"
  },
  {
   "fact": "ActorAttack multi-turn jailbreaks, attacked by the uncensored Dolphin-X1-8B: risk exposure 84.5% at one sample on the HarmBench subset, and 0.0% attack success on both subsets at a sampling budget of 5. Speculation hit rate against the agent's real action was 47.2% on AgentDojo at budget 4 but 8% on ActorAttack.",
   "locator": "v1, Section 4.2.2; Table 2; Section 4.2.6"
  },
  {
   "fact": "On benign utility checks with the judge detector, BFCL-v3 overall accuracy was 45.13 without and 44.88 with SSH, XSTest refusal rate was 0.0% in both, and no benign sample on those two checks triggered an alert; separately, the system-level false-positive rate on AgentDojo, from the limitations section, was 0.2%, and a false positive adds roughly 3 to 7 seconds.",
   "locator": "v1, Section 4.2.3; Table 3; Limitations"
  },
  {
   "fact": "With GPT-4o-mini and GPT-OSS-120B as target agents on AgentDojo, average attack success was 8.4% and 14.9% without SSH and 0.0% and 0.7% with SSH plus the judge detector; the authors call this near-zero.",
   "locator": "v1, Appendix D.1; Table 5"
  },
  {
   "fact": "The 4,700-sample fine-tuning data for injection scenarios were synthesized from 30 attack templates; the paper does not report how closely these resemble the benchmarks' attacks, and it does not test attackers that know SSH is deployed.",
   "locator": "v1, Appendix B.3; Section 4"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-074"
 ],
 "methods": [
  "ai-monitoring",
  "indirect-prompt-injection",
  "injection-classifiers",
  "jailbreaking"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-10-01"
}