Desk/2026-W18

Week of Apr 27 – May 3, 2026

3 records1 status changes on new evidence2 new findings

What changed in what we know

New findings

Capability & gating

May 1, 2026
CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months
CapabilityEvaluation reportUS Center for AI Standards and Innovation, DeepSeek, NIST

NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.

Defense & research

Apr 30, 2026
Microsoft Research red-teams a network of 100+ agents and finds propagation and trust-capture failures
DefensePaperMicrosoft

Microsoft researchers red-teamed an internal platform of over 100 always-on LLM agents that represent different people and interact through forums, messages and a marketplace. They describe four network-level failure modes: self-propagating messages, amplification of false claims, capture of reputation and verification systems, and hard-to-trace flows through unwitting intermediaries. A small share of agents spontaneously adopted protective behaviors that spread through the network.

Policy & standards

May 1, 2026
CISA, ASD's ACSC and international partners publish joint guidance on careful adoption of agentic AI
PolicyGuidanceCISA, NSA, Australian Signals Directorate (ACSC)

CISA and the Australian Signals Directorate, with US and international partners, published guidance on the cybersecurity risks of agentic AI services and recommended mitigations. CISA highlights expanded attack surface, privilege creep, behavioral misalignment and obscure event records as risks, and recommends avoiding broad access, starting with low-risk use cases, and folding agent security into existing risk models.