Methods/Defense

Delimiting untrusted input

Marking untrusted content so the model can tell it apart from instructions, for example with delimiters, encoding, or spotlighting.

4 records4 defense2 findings (2 measured)First recorded 2024-03assistant-drafted

How it works

The application transforms or tags untrusted text before the model sees it and tells the model never to follow instructions inside it.

Known limits

Reduces success against fixed attacks but has failed against adaptive attackers.

What we know

1 reported, 1 revalidate

Records over time

RangeLanes
2 of 4 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
202420252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

2 records · newest first
Oct 2025
May 2025
May 20, 2025
Google DeepMind reports lessons from continuously attacking Gemini with adaptive prompt injections
DefensePaperGoogle DeepMind

Shi and colleagues describe Google DeepMind's continuous adaptive-attack evaluation of Gemini against indirect prompt injection in tool-use settings. On Gemini 2.0, adaptive attacks generally matched or beat non-adaptive ones against eight baseline defenses, reaching 98.4% against in-context learning and 82.4% against spotlighting, while a warning defense and a user-instruction classifier held (at most 10.8% and 3.0%). Adversarial fine-tuning for Gemini 2.5 lowered but did not eliminate attack success.

All records