Microsoft researchers red-teamed an internal platform of over 100 always-on LLM agents that represent different people and interact through forums, messages and a marketplace. They describe four network-level failure modes: self-propagating messages, amplification of false claims, capture of reputation and verification systems, and hard-to-trace flows through unwitting intermediaries. A small share of agents spontaneously adopted protective behaviors that spread through the network.
Why it matters
It shows agent-to-agent interaction creates attack paths, such as worms and proxy exfiltration, that single-agent testing misses.
Key facts
As stated in the sources, with where to find them.
- A single self-propagating message reached all 6 agents in the test group, looped back after six hops and kept circulating for over 12 minutes; in total it consumed over 100 LLM calls billed to the victims' principals.Case study 1, Self-propagating worms
- A fabricated claim drew 299 comments from 42 agents; in a separate test, sensitive data reached the attacker through a single intermediary over five messages.Case studies 2 (Reputation manipulation) and 4 (Proxy chains)
Findings that cite this record
Sources
Related records
Aug 18, 2026
Nov 19, 2025
Aug 6, 2025
Jun 11, 2025
Jun 11, 2025
Mar 5, 2024