Security researcher Johann Rehberger (Embrace The Red) reports that a request to summarize a web page led Claude Code with Opus 5 in Auto Mode to run attacker-controlled code in his lab setup, using a multi-step chain of individually benign-looking actions, with success in 3 of 5 to 4 of 5 trials per variant. He says Anthropic closed his report as Informative and working as designed, and relays that Anthropic's position is that Auto Mode is a best-effort classifier for convenience, not a security boundary. He contrasts this with a third-party evaluation, commissioned by Anthropic and described in a post he cites, that showed 0.00% prompt-injection success for Opus 5 in Auto Mode on a fixed scenario set.
Week of Aug 24–30, 2026
Attacks & incidents
Defense & research
Trail of Bits argues in a blog post that a standard virtual machine should no longer be assumed to contain a sufficiently capable cyber agent, based on one researcher's informal test of a preview of OpenAI's GPT 5.6-Cyber obtained through Patch the Planet. In a capture-the-flag setup on a Linux development machine, the agent was told to escape a QEMU/KVM VM; the author reports three escapes: a known host-kernel bug (whose exploit did not land cleanly), a combination of libslirp bugs on an oldstable distribution, and a chain that included 0-day bugs against freshly rebuilt upstream QEMU. The author recommends rapidly updated distributions, minimal-attack-surface virtualization such as Firecracker (which the agent did not escape), least privilege, logging, monitoring, time limits and a fresh environment per use.