Chronicle/Defense & research

SEC-bench automatically builds real vulnerability tasks and finds agents patch at most 34%

DefenseBenchmarkSignificance assistant-drafted

SEC-bench uses multi-agent scaffolding to construct reproducible vulnerability instances with test environments and validated patches from real projects, at about $0.87 per instance. The authors report that LLM agents reached at most 18.0% on proof-of-concept generation and 34.0% on vulnerability patching.

Why it matters

It offers a cheaper route to fresh vulnerability benchmarks and shows low agent patching rates even with call-stack hints and a build-and-PoC check.

Key facts

As stated in the sources, with where to find them.

  • Dataset construction cost about $0.87 per instance.Abstract
  • On the full 200-instance dataset with Claude 3.7 Sonnet, the best scaffold (OpenHands) reached 18.0% on PoC generation and 34.0% on vulnerability patching; a patch counts if the project builds and the original PoC no longer triggers.Sections 2 and 3.2

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records