Chronicle/Attacks & incidents

Morris II paper demonstrates self-replicating prompts spreading between GenAI email assistants

AttackPaperSignificance assistant-drafted

Cohen, Bitton and Nassi present Morris II, an adversarial self-replicating prompt that propagates through RAG-based GenAI email assistants, causing data exfiltration and further spread. The paper also proposes a detection guardrail and reports its accuracy.

Why it matters

It showed that prompt injection can propagate between connected assistants, a precursor to multi-agent attack concerns.

Key facts

As stated in the sources, with where to find them.

  • The proposed Virtual Donkey guardrail is reported at a true-positive rate of 1.0 with a false-positive rate of 0.015.Abstract
  • Evaluation used GenAI-powered email assistants across context sizes, embedding algorithms and propagation chains.Abstract

Findings that cite this record

Sources

Related records