Researchers at SPAR, Cambridge, APTA AI and CISPA study a class of attack in which adversarial text in a shared artifact is stored in one assistant's memory, reproduced in an artifact it later writes and picked up by another assistant, with no direct agent-to-agent channel. In 36 synthetic workflows on the OpenClaw harness with an attacker-operated upload endpoint, the authors report that a single seed artifact goal-infected 38% to 98% of assistants (32% to 73% fully infected) depending on the model. They report an attacker-service-free variant that was weaker, and that an off-the-shelf classifier at the memory-write step flagged their main template's infected memories, with false positives on legitimate instructions.
It suggests shared files and assistant memory can act together as a spread channel between assistants that never communicate, at least in a simulation with permissive outbound access.
Key facts
As stated in the sources, with where to find them.
- Setup: 36 held-out synthetic human-agent universes (3 to 12 agents, 10 to 20 time steps) on the OpenClaw harness with its default MEMORY.md, each universe run twice per model; one seed artifact; a goal-agnostic attack template optimized separately per target model on 12 development universes and then frozen; an LLM judge labels memories and artifacts.v1, Sections 4.2, 5 and 6.1
- The main experiments use an endpoint-assisted variant in which infected assistants send artifacts to an attacker-operated service that reinserts a clean copy of the attack; the authors say the resulting rates do not measure an assistant's unaided ability to reproduce the attack and that egress restrictions would interrupt this mechanism.v1, Section 4.1; Appendix A (Limitations)
- Fraction of assistants goal-infected by one seed: GPT-5.6 Luna 0.38, Kimi-K2.6 0.47, GPT-OSS-120B 0.85, DeepSeek-V4-Pro 0.98; hop-4 survival 0.24, 0.44, 0.61, 0.76. Full infection (goal plus propagation instruction) was 0.32, 0.37, 0.73, 0.71.v1, Table 1
- In three 30-agent, 60-step universes with three seed artifacts each (one run per universe, described by the authors as descriptive), the attack reached 60% to 80% of agents for GPT-5.6 Luna, 77% to 97% for Kimi-K2.6 and 90% to 100% for GPT-OSS-120B and DeepSeek-V4-Pro; deepest chains were 3 to 8 hops for Luna.v1, Section 6.2; Appendix H
- Attack development effort: 109 jobs for GPT-OSS-120B, 372 for Kimi-K2.6, 932 for DeepSeek-V4-Pro and 4,002 for GPT-5.6 Luna; a search against Grok 4.6 was stopped for cost after 571 jobs without meeting the stopping rule, and the authors say this is not evidence of immunity.v1, Appendix F, Table 6; Section 6.2
- Without the attacker endpoint, the same design on DeepSeek-V4-Flash fully infected 21% of assistants with second-hop survival 0.30, which the authors describe as a lower bound at a modest search budget.v1, Appendix D; Appendix A
- 96% to 100% of infected assistants sent at least one artifact to the attacker endpoint. Re-running the test universes with security-critical goals gave equal or lower spread for every model, for example first-hop survival S(1) of 0.56 for Luna versus 0.80.v1, Appendix I.1, Table 8; Appendix E
- An off-the-shelf classifier asked a generic question flagged 104 of 104 infected memory lists from the main template and all 24 legitimate instruction-bearing memories; a provenance-worded question had 4 of 24 false positives, and memories holding only a false belief without an instruction were mostly missed. The authors ran this test on their own template.v1, Section 7; Appendix J, Table 9
Findings that cite this record
Key questions this bears on
- Can prompt injection against AI agents be reliably defended?Not reliably. Adaptive attackers still beat some 2026 models; bounding what untrusted input can trigger is the best-supported defense.