MonitoringBench shows weak attacks overstate monitor performance.
Week of May 4–10, 2026
What changed in what we know
Attacks & incidents
Microsoft's Defender security research team described two vulnerabilities in its open-source Semantic Kernel agent framework in which attacker-influenced prompts could reach code execution on the host. One involved unsafe string interpolation in an in-memory vector store filter; the other exposed a code-execution plugin function that allowed arbitrary file writes outside its sandbox. Both were fixed in updated SDK releases.
Defense & research
Jotautaitė, Martinez, Matthews and Tracy (two listed as independent, Tracy at Redwood Research) build a semi-automated red-teaming pipeline in the BashArena control setting, where agents covertly pursue side tasks such as firewall backdoors, privilege escalation, credential exfiltration and malware download. MonitoringBench's 2,644 attack trajectories lowered a Claude Opus 4.5 monitor's catch rate at 1% false positives from 94.9% on directly elicited attacks to 60.3% on refined ones.
Policy & standards
The Coalition for Secure AI released a paper on identity and access management for agents from its Secure Design Patterns for Agentic Systems workstream, focused on unique agent credentials and task-limited access. A companion paper on multi-agent systems discusses semantic-layer attacks, intent-based authorization and proposes agent detection and response as a defense category.