Desk/2025-W19

Week of May 5–11, 2025

3 records0 status changes on new evidence1 new findings

New findings

Defense & research

May 7, 2025
UC Santa Cruz study integrates LLM agents into CAGE 4 and finds RL defenders still outperform them
DefensePaperUC Santa Cruz

Researchers led by UC Santa Cruz integrated LLM agents into the CybORG CAGE 4 multi-agent defence environment and proposed a communication protocol for mixed LLM and RL teams. In their runs an all-RL team scored far better reward than an all-LLM (GPT-4o-mini) team and acted about 104 times faster, though the authors highlight LLM explainability and note the environment was designed for RL agents.

May 6, 2025
Meta releases LlamaFirewall guardrails with PromptGuard 2 and AlignmentCheck for agents
DefenseTool releaseMeta

Meta open-sources LlamaFirewall, combining PromptGuard 2 (a jailbreak and injection detector), AlignmentCheck (a chain-of-thought auditor for goal hijacking) and CodeShield (static analysis of generated code). On AgentDojo, Meta reports that the combination cut attack success from 17.63% to 1.75% while utility fell from 47.73% to 42.68%.

Policy & standards

May 7, 2025
UK NCSC judges AI-assisted vulnerability research is the most significant AI cyber development to 2027
PolicyGuidanceUK National Cyber Security Centre

The NCSC's second assessment judges that AI will almost certainly make elements of intrusion more effective through 2027, with AI-assisted vulnerability research and exploit development the most significant development. It warns that the window between disclosure and exploitation, already days, will shrink further, and judges fully automated end-to-end advanced attacks unlikely before 2027.