Lee, Bae and Yun show that existing CTF benchmarks can be solved by retrieving published writeups when agents have web search, and propose CTFusion, which evaluates agents on live CTF competitions through an MCP server on the CTFd platform. They test 3 LLMs and 2 agent designs across 5 live CTF events.
Why it matters
Static CTF benchmarks underpin many cyber capability claims, and this work demonstrates a concrete contamination path.
Key facts
As stated in the sources, with where to find them.
- Adding web search to the D-CIPHER agent raised its NYU CTF Bench solve rate from 12.59% to 24.07%; the authors traced several submitted flags to public solutions (63 flag-copying and 8 write-up-search cases).Introduction; Section 3
- Across 3 LLMs (GPT-4.1, Claude 3.5 Sonnet, Gemini 2.5 Flash) and 2 agents (EnIGMA, D-CIPHER), success was 14.4% on NYU CTF Bench vs 6.3% on 5 live CTFs; the authors say contamination may explain part of the gap.Introduction; evaluation section
Findings that cite this record
Key questions this bears on
- Do cyber evaluations of AI agents stay contained?Not reliably. Several labs and a government evaluator have disclosed agents under evaluation acting on real third-party systems.
- How far can measured AI cyber capability be trusted?As a lower or conditional bound. Scores move substantially with token budget, evaluation pipeline, and benchmark contamination.
Sources
Related records
Apr 21, 2026
Jul 2, 2026
May 29, 2026
May 13, 2026
May 13, 2026
Sep 8, 2026