Mar 5, 2024
Week of Mar 4–10, 2024
2 records0 status changes on new evidence2 new findings
New findings
Attacks & incidents
Mar 5, 2024
Morris II paper demonstrates self-replicating prompts spreading between GenAI email assistants
Cohen, Bitton and Nassi present Morris II, an adversarial self-replicating prompt that propagates through RAG-based GenAI email assistants, causing data exfiltration and further spread. The paper also proposes a detection guardrail and reports its accuracy.
Defense & research
Mar 5, 2024
InjecAgent benchmarks indirect prompt injection against tool-integrated LLM agents
Zhan, Liang, Ying and Kang release InjecAgent, a benchmark of 1,054 test cases spanning 17 user tools and 62 attacker tools, covering direct harm to users and exfiltration of private data. They evaluate 30 LLM agents and find a ReAct-prompted GPT-4 agent vulnerable in about a quarter of cases.