NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.
Week of Sep 29 – Oct 5, 2025
Capability & gating
Defense & research
Anthropic reports that a small team focused Claude Sonnet 4.5 training on finding and patching vulnerabilities and on testing simulated security infrastructure, while avoiding enhancements that clearly favour offence. It reports Sonnet 4.5 results on Cybench and CyberGym, a preliminary patching study in which 15% of patches were judged semantically equivalent to human references, and invites work on SOC and SIEM automation.
Policy & standards
MITRE ATLAS version 5.0.0 added a set of techniques for attacks on AI agents, including agent context poisoning of memory and threads, modifying agent configuration, credential theft from agent configuration, and exfiltration via agent tool invocation, and renamed LLM Plugin Compromise to AI Agent Tool Invocation. Version 5.1.0 (November 6, 2025) added agent-specific mitigations such as tool permission configuration and human-in-the-loop for agent actions.
California's Transparency in Frontier AI Act (SB 53) requires large frontier developers to publish frontier AI frameworks addressing catastrophic risk, model weight cybersecurity and incident response, and to report critical safety incidents to the Office of Emergency Services. Its catastrophic risk definition includes a model engaging, with no meaningful human oversight, in conduct that is a cyberattack, where a single incident causes death or serious injury to more than 50 people or more than $1 billion in property damage.