Desk/2025-W40

Week of Sep 29 – Oct 5, 2025

4 records0 status changes on new evidence0 new findings

Capability & gating

Sep 30, 2025
CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack
CapabilityEvaluation reportUS Center for AI Standards and Innovation, NIST, DeepSeek

NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.

Defense & research

Oct 3, 2025
Anthropic says it trained Claude Sonnet 4.5 for defensive vulnerability finding and patching
DefensePaperAnthropic, HackerOne, CrowdStrike

Anthropic reports that a small team focused Claude Sonnet 4.5 training on finding and patching vulnerabilities and on testing simulated security infrastructure, while avoiding enhancements that clearly favour offence. It reports Sonnet 4.5 results on Cybench and CyberGym, a preliminary patching study in which 15% of patches were judged semantically equivalent to human references, and invites work on SOC and SIEM automation.

Policy & standards

Sep 30, 2025
MITRE ATLAS 5.0 adds AI agent techniques such as context poisoning and exfiltration via tool invocation
PolicyStandardMITRE

MITRE ATLAS version 5.0.0 added a set of techniques for attacks on AI agents, including agent context poisoning of memory and threads, modifying agent configuration, credential theft from agent configuration, and exfiltration via agent tool invocation, and renamed LLM Plugin Compromise to AI Agent Tool Invocation. Version 5.1.0 (November 6, 2025) added agent-specific mitigations such as tool permission configuration and human-in-the-loop for agent actions.

Sep 29, 2025
California SB 53 requires frontier AI frameworks covering autonomous cyberattack risk and incident reporting
PolicyRegulationState of California

California's Transparency in Frontier AI Act (SB 53) requires large frontier developers to publish frontier AI frameworks addressing catastrophic risk, model weight cybersecurity and incident response, and to report critical safety incidents to the Office of Emergency Services. Its catastrophic risk definition includes a model engaging, with no meaningful human oversight, in conduct that is a cyberattack, where a single incident causes death or serious injury to more than 50 people or more than $1 billion in property damage.