Desk/2025-W06

Week of Feb 3–9, 2025

2 records0 status changes on new evidence1 new findings

New findings

Defense & research

Feb 3, 2025
Anthropic introduces Constitutional Classifiers against universal jailbreaks
DefensePaperAnthropic

Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.

Policy & standards

Feb 6, 2025
Cloud Security Alliance publishes MAESTRO seven-layer threat modeling framework for agentic AI
PolicyFrameworkCloud Security Alliance

The Cloud Security Alliance published MAESTRO (Multi-Agent Environment, Security, Threat, Risk, and Outcome), a threat modeling framework for agentic AI authored by Ken Huang. It organizes analysis into seven layers from foundation models to the agent ecosystem and highlights agent-specific threats such as goal manipulation, agent impersonation and collusion between agents.