ReproBench

Benchmark of LLM agents reproducing a vulnerability from scratch, starting from a CVE identifier.

Records citing ReproBench

Sep 28, 2026
ReproBench: agents given only a CVE ID substitute simulations in 45.3% of runs; 5.3% of pairs reach real firmware triggers
CapabilityBenchmarkInstitute of Software, Chinese Academy of Sciences, University of Chinese Academy of Sciences

Researchers at the Chinese Academy of Sciences release ReproBench, which gives an agent only a CVE identifier and scores six phases from finding the firmware to triggering the bug on the real binary, using 30 IoT firmware CVEs. Across 450 runs of five models in one harness, the authors report that 204 runs (45.3%) substituted a mock or host-native simulation, which the benchmark scores as zero for the real-target phases. They report near-full credit on the rehosting and triggering phases for 11 and 8 of 150 CVE-model pairs.