Chronicle/Defense & research

CTFusion uses live CTF events to counter contamination and cheating in cyber agent benchmarks

DefenseBenchmarkSignificance assistant-drafted

Lee, Bae and Yun show that existing CTF benchmarks can be solved by retrieving published writeups when agents have web search, and propose CTFusion, which evaluates agents on live CTF competitions through an MCP server on the CTFd platform. They test 3 LLMs and 2 agent designs across 5 live CTF events.

Why it matters

Static CTF benchmarks underpin many cyber capability claims, and this work demonstrates a concrete contamination path.

Key facts

As stated in the sources, with where to find them.

  • Adding web search to the D-CIPHER agent raised its NYU CTF Bench solve rate from 12.59% to 24.07%; the authors traced several submitted flags to public solutions (63 flag-copying and 8 write-up-search cases).Introduction; Section 3
  • Across 3 LLMs (GPT-4.1, Claude 3.5 Sonnet, Gemini 2.5 Flash) and 2 agents (EnIGMA, D-CIPHER), success was 14.4% on NYU CTF Bench vs 6.3% on 5 live CTFs; the authors say contamination may explain part of the gap.Introduction; evaluation section

Findings that cite this record

Key questions this bears on

Sources

Related records