Week of Feb 3–9, 2025
New findings
Defense & research
Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.
Policy & standards
The Cloud Security Alliance published MAESTRO (Multi-Agent Environment, Security, Threat, Risk, and Outcome), a threat modeling framework for agentic AI authored by Ken Huang. It organizes analysis into seven layers from foundation models to the agent ecosystem and highlights agent-specific threats such as goal manipulation, agent impersonation and collusion between agents.