CyberPersistBench

Benchmark of LLM-based attackers on installing and persisting after a compromise.

Records citing CyberPersistBench

Sep 29, 2026
CyberPersistBench scores agent post-compromise persistence: 27.6% to 42.4% on 203 tasks, 5.5% to 13.3% under native defenses
CapabilityBenchmarkShanghai Artificial Intelligence Laboratory

Researchers at Shanghai AI Laboratory release CyberPersistBench, which starts agents from a restricted foothold and scores whether they keep durable access after credentials are revoked and services or hosts are disrupted. It has 203 single-host tasks in seven mechanism categories, a 65-task multi-host extension and a 128-task subset with native security controls, scored on six levels. The authors report pass@3 success of 27.6% to 42.4% for five models on the common scaffold and 5.5% to 13.3% with defenses enabled.