HKUST researchers argue that agentic vulnerability discovery forms many hypotheses but can verify only some within a fixed budget, so verification effort is something a defender can steer. Their RedHerring system inserts decoy code paths that look like real CVE-style bugs but are unreachable behind a predicate the defender can certify with private information. In 70 OSS-Fuzz instances, the authors report that five open-weight models run in Claude Code confirmed 38.7% to 60.4% fewer real crashes, with 30.6% to 51.5% of completion tokens spent on decoys.
It tests a defense aimed at attackers' AI vulnerability-search budgets rather than at individual bugs, but only on open-weight models and single runs, so it does not show how frontier agents would respond.
Key facts
As stated in the sources, with where to find them.
- In 300 Qwen3.8-Flash trajectories annotated by an LLM, verification took 73.8% of wall-clock time (95% CI 71.7% to 75.5%) and one verification cost 11.1 times as much time as forming one hypothesis; agents formed 9,551 hypotheses, started verifying 7,301 and abandoned 4,576 (62.7%). Ambiguous steps were assigned to verification, so the authors call 73.8% an upper bound.v1, Section 3.1; Appendix B
- Evaluation: 70 instances from 33 OSS-Fuzz projects, Baseline versus five inserted decoys, five open-weight models (Qwen3.8-Flash, Qwen3.8-Max, GLM-5.3, DeepSeek-V4-Pro, Kimi-K3) in Claude Code with a 3-hour, 300-round limit; each configuration run once. The metric is distinct confirmed crash signatures replayed on a local verification server, used as a proxy for vulnerabilities.v1, Section 5.1; Appendix A
- Confirmed real vulnerabilities, Baseline to RedHerring, summed over 70 instances: Qwen3.8-Flash 149 to 59 (-60.4%); Qwen3.8-Max 90 to 40 (-55.6%); Kimi-K3 98 to 54 (-44.9%); DeepSeek-V4-Pro 84 to 50 (-40.5%); GLM-5.3 124 to 76 (-38.7%).v1, Figure 4a
- Decoys took 30.6% to 51.5% of completion tokens and an estimated 32.5% to 49.9% of runtime, with shares labeled by Qwen3.8-Flash plus symbol matching; the authors report a Pearson correlation of 0.99 between token share and reduction across the five models.v1, Figure 4b-c; Section 5.3; Appendix A
- When the prompt told Qwen3.8-Flash that decoys might be present, RedHerring still cut confirmed vulnerabilities by 37.2% against an informed Baseline (121 to 76), with decoys taking about 30% of completion tokens; the notice also lowered the Baseline count. Only this one model was tested in the informed setting.v1, Section 5.4, Figure 6
- Ablations on Qwen3.8-Flash: removing the vulnerability chain gives a 25.5% reduction, replacing the false bridge with a simple always-false condition 32.9%, and harmless code behind the same condition 14.1%; the reduction grows as decoys per instance rise from one to five.v1, Section 5.5, Figure 7
- Defender cost: 30 hours of one-time offline preparation for 70 instances, 69.3 minutes on average to integrate five decoys, 13.8% larger source, under 1% runtime overhead on native test suites, all tests passing, and all 350 inserted false bridges passing the safety checks.v1, Section 5.6
- The authors state that Claude and GPT models were excluded because of safety-alignment refusals, that runs were single per configuration, and that the effort shares depend on an annotation model.v1, Appendix A (Limitations)
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.