Anthropic describes three defenses for browser use: reinforcement learning on injected web content, classifiers that scan untrusted content, and human red-teaming including external arena-style challenges. Against an internal adaptive Best-of-N attacker with 100 attempts per environment, its chart shows Claude Opus 4.5 with the improved Claude for Chrome safeguards at 1.4% attack success (8.0% for the model alone), which the post rounds to 1%. Anthropic says this still represents meaningful risk.
Why it matters
It gives a lab-reported adaptive-attacker metric for a deployed browser agent and explicitly treats residual 1% ASR as unsolved.
Key facts
As stated in the sources, with where to find them.
- Claude Opus 4.5 in Claude for Chrome, internal adaptive Best-of-N attacker: 8.0% ASR model only, 7.6% with the old safeguards, 1.4% with the improved safeguards (N=2,782 attacks encountered); the post's text rounds this to 1%.Chart 'Claude for Chrome Prompt Injection Robustness' and caption
- With the improved safeguards, the other models tested scored 2.0% (Sonnet 4), 1.9% (Sonnet 4.5) and 2.4% (Haiku 4.5).Same chart
- The attacker gets 100 attempts per environment; ASR is computed as a share of the attacks each model actually encountered.Chart caption
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Feb 5, 2026
Jan 9, 2026
Mar 16, 2026
Sep 30, 2025
Oct 10, 2025
Jan 17, 2025