Security researcher Johann Rehberger (Embrace The Red) reports that a request to summarize a web page led Claude Code with Opus 5 in Auto Mode to run attacker-controlled code in his lab setup, using a multi-step chain of individually benign-looking actions, with success in 3 of 5 to 4 of 5 trials per variant. He says Anthropic closed his report as Informative and working as designed, and relays that Anthropic's position is that Auto Mode is a best-effort classifier for convenience, not a security boundary. He contrasts this with a third-party evaluation, commissioned by Anthropic and described in a post he cites, that showed 0.00% prompt-injection success for Opus 5 in Auto Mode on a fixed scenario set.
It is a researcher's small-sample targeted test of a shipped classifier-based approval mode, and the vendor response he reports places the security boundary at OS isolation and network egress control rather than at the classifier.
Key facts
As stated in the sources, with where to find them.
- The post says Auto Mode replaces human approval prompts with a safety classifier and has been the default starting mode in Claude Code since mid-August.Auto Mode Is Now the Default in Claude Code
- Bypass class: indirect prompt injection through attacker-controlled web content, which steers the agent, through actions that each look benign, into running code from an attacker-controlled directory; the agent's refusal to run an obviously malicious program did not prevent this. In the author's lab setup the code reached a command-and-control server he controlled.In A Nutshell; Walkthrough
- Reported success rates, small samples and not a universal measure: 3/5 (60%) for a variant with a remote stager and callback, 3/5 (60%) and 4/5 (80%) for two variants that launch a nested Claude Code instance. The author says rates improved as payloads were iterated with Codex and that results are representative of a motivated attacker but not comprehensive.Reliability and Attack Success Rate (ASR)
- In some runs the agent avoided compromise, for example by analyzing the archive statically or running its helper from a safe directory; in others it noticed the compromise afterward, and in a few runs Auto Mode denied the agent's command to stop the malicious process.Claude Realizes Attacks Too Late; Auto Mode Blocks Cleanup; Refusals
- Disclosure: the author says he emailed Anthropic's model bug bounty address without a response, then used its security reporting channel, and Anthropic closed the report as Informative and working as designed.Disclosure
- The author relays Anthropic's position as: Auto Mode is a convenience feature with a best-effort classifier, not a security guarantee, and the real boundary is OS isolation and network egress control. This is the researcher's account, not an Anthropic statement.Disclosure
- Per the post, a Trajectory Labs evaluation commissioned by Anthropic ran 72 indirect prompt-injection scenarios ten times each and showed 0.00% attack success for Opus 5 in Auto Mode; the author says his chain was not in that set. The evaluation itself was not reviewed for this record.Auto Mode Is Now the Default; The 0.00% Marketing Problem
- The author recommends running unattended coding agents in a container, VM or OS sandbox, restricting network egress, monitoring agents, and not exposing home directories or credentials.Mitigation: Sandboxing - Not Optional
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.
- Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.