Chronicle/Defense & research

UC Santa Cruz study integrates LLM agents into CAGE 4 and finds RL defenders still outperform them

DefensePaperSignificance assistant-drafted

Researchers led by UC Santa Cruz integrated LLM agents into the CybORG CAGE 4 multi-agent defence environment and proposed a communication protocol for mixed LLM and RL teams. In their runs an all-RL team scored far better reward than an all-LLM (GPT-4o-mini) team and acted about 104 times faster, though the authors highlight LLM explainability and note the environment was designed for RL agents.

Why it matters

The authors describe it as the first study of LLM agents in a multi-agent autonomous cyber defense environment, and it cautions against assuming LLMs beat trained RL policies.

Key facts

As stated in the sources, with where to find them.

  • Mean reward: all-RL team (KEEP) -493 (sd 95.9) vs all-LLM team with GPT-4o-mini -2547.2 (sd 498.8).Section IV-A, Figure 5
  • RL agents were about 104.1 times faster at action selection; all-RL runs averaged 45.2 s vs 4704.6 s for all-LLM (GPT-4o-mini).Section IV-A
  • Models tested: GPT-4o-mini, o3-mini, o1-mini and DeepSeek-V3; 2 episodes of 500 steps per scenario.Experimental setup
  • Presented at the 2025 IEEE CAI Workshop on Adaptive Cyber Defense.arXiv comments

Findings that cite this record

Sources

Related records