Researchers led by UC Santa Cruz integrated LLM agents into the CybORG CAGE 4 multi-agent defence environment and proposed a communication protocol for mixed LLM and RL teams. In their runs an all-RL team scored far better reward than an all-LLM (GPT-4o-mini) team and acted about 104 times faster, though the authors highlight LLM explainability and note the environment was designed for RL agents.
Why it matters
The authors describe it as the first study of LLM agents in a multi-agent autonomous cyber defense environment, and it cautions against assuming LLMs beat trained RL policies.
Key facts
As stated in the sources, with where to find them.
- Mean reward: all-RL team (KEEP) -493 (sd 95.9) vs all-LLM team with GPT-4o-mini -2547.2 (sd 498.8).Section IV-A, Figure 5
- RL agents were about 104.1 times faster at action selection; all-RL runs averaged 45.2 s vs 4704.6 s for all-LLM (GPT-4o-mini).Section IV-A
- Models tested: GPT-4o-mini, o3-mini, o1-mini and DeepSeek-V3; 2 episodes of 500 steps per scenario.Experimental setup
- Presented at the 2025 IEEE CAI Workshop on Adaptive Cyber Defense.arXiv comments
Findings that cite this record
Sources
Related records
Feb 20, 2024
Sep 4, 2026
Aug 4, 2026
Jul 21, 2026
Jun 27, 2025
May 21, 2025