PatchBench

Benchmark that verifies AI-generated patches with several independent checks, not only the original proof of concept.

Records citing PatchBench

Sep 3, 2026
PatchBench finds PoC-only checks inflate AI patching success 1.83x and 25% of patches look memorized
DefenseBenchmarkUniversity of Maryland AI Security Lab

PatchBench, from the University of Maryland's AI Security Lab, evaluates 11 patching agents, including the top three AIxCC systems, on 213 C/C++ tasks whose true fixes lie outside the crash stack, using vulnerability transplant and code mutation to limit memorization. It finds that accepting a patch because the original proof-of-concept no longer crashes inflates solve rates by 1.83x on average, and that about 25% of agent patches closely resemble historical developer fixes.