Announcing a limited pilot of Claude in Chrome, Anthropic reports red-teaming with 123 test cases across 29 attack scenarios. Attack success in autonomous mode was 23.6% without new mitigations and 11.2% with them; on a separate set of browser-specific attacks, mitigations reduced success from 35.7% to 0%.
Why it matters
Anthropic published a non-trivial residual prompt injection rate for a browser agent it was piloting with users, not only the improvement from its mitigations.
Key facts
As stated in the sources, with where to find them.
- 123 test cases, 29 attack scenarios: 23.6% ASR without mitigations, 11.2% with mitigations in autonomous mode.Safety section
- Browser-specific attack set (e.g., hidden form fields, URL and tab-title injections): 35.7% to 0% with mitigations.Safety section
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Defenses reduce injection but none has eliminated it; limiting what untrusted input can trigger is the most defensible approach.
Sources
Related records
Aug 6, 2025
Feb 5, 2026
Jan 20, 2026
Oct 20, 2025
Jan 28, 2026
Oct 31, 2025