{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/redherring-decoys-divert-agent-verification-2026/",
 "asOf": "2026-10-01",
 "id": "redherring-decoys-divert-agent-verification-2026",
 "date": "2026-09-28",
 "datePrecision": "day",
 "title": "RedHerring decoys cut real vulnerabilities found by five open-weight agents by 38.7% to 60.4% at matched budgets",
 "lane": "defense",
 "kind": "paper",
 "summary": "HKUST researchers argue that agentic vulnerability discovery forms many hypotheses but can verify only some within a fixed budget, so verification effort is something a defender can steer. Their RedHerring system inserts decoy code paths that look like real CVE-style bugs but are unreachable behind a predicate the defender can certify with private information. In 70 OSS-Fuzz instances, the authors report that five open-weight models run in Claude Code confirmed 38.7% to 60.4% fewer real crashes, with 30.6% to 51.5% of completion tokens spent on decoys.",
 "whyItMatters": "It tests a defense aimed at attackers' AI vulnerability-search budgets rather than at individual bugs, but only on open-weight models and single runs, so it does not show how frontier agents would respond.",
 "actors": [
  "hong-kong-university-of-science-and-technology"
 ],
 "topics": [
  "vulnerability-discovery",
  "ai-enabled-intrusion"
 ],
 "atlas": [
  "untrusted-content"
 ],
 "artifacts": [
  "oss-fuzz",
  "deepseek",
  "kimi",
  "glm"
 ],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.35909",
   "publisher": "arXiv",
   "title": "Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery",
   "date": "2026-09-28",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "In 300 Qwen3.8-Flash trajectories annotated by an LLM, verification took 73.8% of wall-clock time (95% CI 71.7% to 75.5%) and one verification cost 11.1 times as much time as forming one hypothesis; agents formed 9,551 hypotheses, started verifying 7,301 and abandoned 4,576 (62.7%). Ambiguous steps were assigned to verification, so the authors call 73.8% an upper bound.",
   "locator": "v1, Section 3.1; Appendix B"
  },
  {
   "fact": "Evaluation: 70 instances from 33 OSS-Fuzz projects, Baseline versus five inserted decoys, five open-weight models (Qwen3.8-Flash, Qwen3.8-Max, GLM-5.3, DeepSeek-V4-Pro, Kimi-K3) in Claude Code with a 3-hour, 300-round limit; each configuration run once. The metric is distinct confirmed crash signatures replayed on a local verification server, used as a proxy for vulnerabilities.",
   "locator": "v1, Section 5.1; Appendix A"
  },
  {
   "fact": "Confirmed real vulnerabilities, Baseline to RedHerring, summed over 70 instances: Qwen3.8-Flash 149 to 59 (-60.4%); Qwen3.8-Max 90 to 40 (-55.6%); Kimi-K3 98 to 54 (-44.9%); DeepSeek-V4-Pro 84 to 50 (-40.5%); GLM-5.3 124 to 76 (-38.7%).",
   "locator": "v1, Figure 4a"
  },
  {
   "fact": "Decoys took 30.6% to 51.5% of completion tokens and an estimated 32.5% to 49.9% of runtime, with shares labeled by Qwen3.8-Flash plus symbol matching; the authors report a Pearson correlation of 0.99 between token share and reduction across the five models.",
   "locator": "v1, Figure 4b-c; Section 5.3; Appendix A"
  },
  {
   "fact": "When the prompt told Qwen3.8-Flash that decoys might be present, RedHerring still cut confirmed vulnerabilities by 37.2% against an informed Baseline (121 to 76), with decoys taking about 30% of completion tokens; the notice also lowered the Baseline count. Only this one model was tested in the informed setting.",
   "locator": "v1, Section 5.4, Figure 6"
  },
  {
   "fact": "Ablations on Qwen3.8-Flash: removing the vulnerability chain gives a 25.5% reduction, replacing the false bridge with a simple always-false condition 32.9%, and harmless code behind the same condition 14.1%; the reduction grows as decoys per instance rise from one to five.",
   "locator": "v1, Section 5.5, Figure 7"
  },
  {
   "fact": "Defender cost: 30 hours of one-time offline preparation for 70 instances, 69.3 minutes on average to integrate five decoys, 13.8% larger source, under 1% runtime overhead on native test suites, all tests passing, and all 350 inserted false bridges passing the safety checks.",
   "locator": "v1, Section 5.6"
  },
  {
   "fact": "The authors state that Claude and GPT models were excluded because of safety-alignment refusals, that runs were single per configuration, and that the effort shares depend on an annotation model.",
   "locator": "v1, Appendix A (Limitations)"
  }
 ],
 "significance": 3,
 "fideQuestions": [],
 "methods": [
  "ai-vulnerability-discovery"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}