An author project page describes DeltaCert-Agent, which maps configuration changes in tool-using LLM agents to affected security claims and reruns only scoped tests plus sentinel checks, escalating to full recertification when impact cannot be bounded. The author reports 75.02% regression-detection recall versus 55.01% for equal-budget random selection while running 61.35% fewer tests, using four small locally hosted models.
Why it matters
Continuous agent changes make full security re-evaluation costly, and this work tests a cheaper recertification strategy.
Key facts
As stated in the sources, with where to find them.
- 75.02% regression-detection recall vs 55.01% for equal-budget random selection; 61.35% fewer tests executed on average.Project page, results
- 31,396 evidence rows across Qwen3, Gemma3, Llama 3.2 and Phi-4 Mini over five repetitions.Project page, evaluation setup
Findings that cite this record
No tracked finding cites this record yet.
Sources
Related records
Apr 29, 2025
Sep 11, 2026
Sep 8, 2026
Sep 3, 2026
Aug 4, 2026
Jul 30, 2026