Scope: what this does not show
One contamination measurement (CTFusion, on NYU CTF Bench) and one position paper that argues, without new measurements, that static benchmarks go stale.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
May 12, 2026
CTFusion uses live CTF events to counter contamination and cheating in cyber agent benchmarks
On NYU CTF Bench, web search raised the D-CIPHER agent's solve rate from 12.59% to 24.07%, with 71 logged cases of copied flags or retrieved writeups.
May 21, 2026
Position paper argues agent security benchmarks suffer from hackable environments, staleness and runtime noise
Argues, without new measurements, that static benchmarks go stale as vulnerabilities are patched and fixes are memorized.
Key questions that rely on this finding
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Status history
- 2026-05-12ReportedCTFusion shows web-searching agents inflate public CTF scores. · record
- 2026-05-21CorroboratedA separate paper argues static cyber benchmarks decay. · record
- 2026-09-25ReportedcorrectionThe 2026-05-21 paper argues that benchmarks go stale but does not measure or independently test contamination, so the measured part of this claim rests on CTFusion alone. · record