Week of Feb 23 – Mar 1, 2026
New findings
Attacks & incidents
Check Point Research found that a cloned repository's Claude Code configuration could run hooks, start MCP servers before the user approved them, and redirect API traffic so the user's Anthropic API key was sent to an attacker (CVE-2025-59536, CVE-2026-21852). Anthropic fixed the issues between August and December 2025 by deferring execution and API calls until after the trust dialog.
Defense & research
Researchers at the UK AI Security Institute and the University of Oxford built SandboxEscapeBench, 18 nested-sandbox capture-the-flag scenarios covering orchestration, runtime and kernel weaknesses rated difficulty 1 to 5. In the March 2026 version, which tested nine models, GPT-5 and Claude Opus 4.5 succeeded on about half of all runs, difficulty 1-2 scenarios were solved at high rates, no model succeeded at difficulty 4 or 5, and success rose roughly log-linearly with token budget. An August 2026 revision reports that newer models did markedly better, with Claude Mythos Preview saturating the public benchmark.
Policy & standards
Anthropic replaced its Responsible Scaling Policy with version 3.0, introducing Frontier Safety Roadmaps and Risk Reports and restating capability thresholds alongside recommended industry-wide mitigations. The published v3.0 policy document does not mention cyber capability; cyber safeguards for later models (Mythos, Fable 5) were described in separate announcements. Versions 3.1 through 3.4 followed between April and July 2026.