Chronicle/Defense & research

Anthropic reports 1.4% prompt injection success for Claude Opus 4.5 with improved Chrome extension safeguards

DefenseEvaluation reportSignificance assistant-drafted

Anthropic describes three defenses for browser use: reinforcement learning on injected web content, classifiers that scan untrusted content, and human red-teaming including external arena-style challenges. Against an internal adaptive Best-of-N attacker with 100 attempts per environment, its chart shows Claude Opus 4.5 with the improved Claude for Chrome safeguards at 1.4% attack success (8.0% for the model alone), which the post rounds to 1%. Anthropic says this still represents meaningful risk.

Why it matters

It gives a lab-reported adaptive-attacker metric for a deployed browser agent and explicitly treats residual 1% ASR as unsolved.

Key facts

As stated in the sources, with where to find them.

  • Claude Opus 4.5 in Claude for Chrome, internal adaptive Best-of-N attacker: 8.0% ASR model only, 7.6% with the old safeguards, 1.4% with the improved safeguards (N=2,782 attacks encountered); the post's text rounds this to 1%.Chart 'Claude for Chrome Prompt Injection Robustness' and caption
  • With the improved safeguards, the other models tested scored 2.0% (Sonnet 4), 1.9% (Sonnet 4.5) and 2.4% (Haiku 4.5).Same chart
  • The attacker gets 100 attempts per environment; ASR is computed as a share of the attacks each model actually encountered.Chart caption

Findings that cite this record

Key questions this bears on

Sources

Related records