A paper in Applied AI Letters by QinetiQ researchers sets out a test and evaluation process for cyber defence agents covering performance, effectiveness, resilience and generalisability, and demonstrates its low-fidelity stage on CAGE Challenge 2 RL agents in CybORG. It introduces Measures of Effectiveness tailored to cyber defence alongside RL reward and tests agents under environment perturbations not seen in training.
Why it matters
It proposes defence-specific effectiveness metrics and robustness tests to complement RL reward when judging whether a defensive agent can be trusted.
Key facts
As stated in the sources, with where to find them.
- The process evaluates performance, effectiveness, resilience and generalisability in low- and high-fidelity environments; the paper demonstrates the low-fidelity stage on CAGE Challenge 2 agents.Abstract
- Agents are evaluated against perturbed conditions to test robustness to scenarios not seen during training.Abstract
Findings that cite this record
No tracked finding cites this record yet.
Sources
Related records
Aug 30, 2025
June 2023
Aug 8, 2025
May 7, 2025
Apr 29, 2025
Sep 5, 2025