Do production prompt-injection defenses survive independent adaptive attack?

How do the 2026 production defenses that labs report as robust perform under independent adaptive attack?

assistant-draftedPrompt injection

Signals

Rests on one source

Current robustness figures for browser and computer-use agents are reported by the labs that built them.

Why it matters

Agents with browsing and tool access are shipping to users. Self-reported rates have a poor track record once attackers adapt.

Hypothesis

Adaptive attack success is several times the rates labs report, and differs more across products than across models.

A first study

Run a pre-registered adaptive-attack protocol against publicly available agent products on synthetic accounts and data, with coordinated disclosure.

Controls it would need

Synthetic data only; attack budgets fixed in advance; disclosure to vendors before publication.

What it could and could not claim

Would describe tested product versions at a point in time; would not publish working attacks.