Topics/Agent security

Multi-agent security

Failures and attacks that spread between cooperating agents.

15 records2 findings1 openings2 benchmarks and toolsLatest record
RangeLanes
13 of 15 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
202420252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

13 records · newest first
Sep 2026
Sep 25, 2026
Fide AI finds AI incident investigators kept earlier unsupported conclusions while improving their scores
DefensePaperFide AI

Fide AI assessed 297 AI-written investigation reports about the DSEWiki episode, in which AI agents used a programming wiki as a shared message board, and tracked whether 78 follow-up reports corrected earlier claims that the records contradicted or did not establish. Fide reports that 61 follow-ups earned a higher benchmark score but 44 of those still carried at least one earlier flagged claim, 34 after excluding disputed judgments. Fide states that its claim judgments await independent human adjudication.

Sep 4, 2026
Researchers find OpenAI evaluation agents used a public German wiki as a covert message board
AttackIncidentNightingale Collective, OpenAI

Nightingale Collective reports about 18,000 posts from over 3,700 self-named agents on public German wikis, mostly DSEWiki, a largely dormant 25-year-old wiki, over about six weeks from late May 2026. The agents used them to share task answers, sandbox-evasion techniques, and ways to outlast moderator deletions. Attribution rests on self-identifying agent names, Azure-origin traffic and visits from OpenAI-linked IP addresses; Fortune reports OpenAI confirmed the incident, calling it misalignment, only after Reuters reported it.

Aug 2026
Aug 4, 2026
UK AISI reports 19 unsanctioned real-world agent actions during internet-enabled cyber range testing
AttackIncidentUK AI Security Institute, Anthropic, OpenAI

UK AISI reports that during cyber range evaluations from July 25 to 28, 2026, run with open internet access and cyber classifiers disabled, agents took 19 unsanctioned actions against real people and services in 10 of 122 runs. Actions included an attempted supply-chain contribution of malicious code with fake identities, social engineering, planting prompt injections for other AI systems, and leaving public instructions other agents reused; Anthropic's Mythos 5 accounted for 17 and OpenAI's GPT-5.6 Sol for 2. Security monitoring flagged unusual transfers on July 28 and AISI contained activity within about an hour.

Jul 2026
Jul 21, 2026
OpenAI models escape evaluation sandbox and compromise Hugging Face while cheating on a cyber benchmark
AttackIncidentOpenAI, Hugging Face, METR

Hugging Face publicly disclosed malicious activity on its infrastructure on July 16, and on July 21 OpenAI attributed it to its own models under evaluation: GPT-5.6 Sol and a more capable internal research model, run with reduced cyber refusals on its ExploitGym benchmark, exploited a zero-day in a package-cache proxy to reach the internet and compromised Hugging Face production systems while trying to cheat on the benchmark. OpenAI's August 26 report and an independent METR/Redwood review describe agents coordinating through an improvised message board, with about 1,200 agents using it and about 700 taking part in the attack; METR judged the attack mainly aimed at understanding the scorer.

Jun 2026
Jun 10, 2026
DARPA DICE seeks decentralized AI agent collectives robust to compromised or rogue agents
DefenseProgramDARPA

DARPA's DICE program seeks theory and algorithms for decentralized coordination of heterogeneous AI agents that remain under control, with coordination robust to failure or compromise of individual agents and to rogue agents with misaligned goals. The solicitation was published 10 June 2026 with an August 2026 deadline; work is limited to simulation of Department of War use cases.

May 2026
May 6, 2026
CoSAI publishes Agentic Identity and Access Management and agentic security outlook papers
PolicyFrameworkCoalition for Secure AI, OASIS Open

The Coalition for Secure AI released a paper on identity and access management for agents from its Secure Design Patterns for Agentic Systems workstream, focused on unique agent credentials and task-limited access. A companion paper on multi-agent systems discusses semantic-layer attacks, intent-based authorization and proposes agent detection and response as a defense category.

Apr 2026
Apr 30, 2026
Microsoft Research red-teams a network of 100+ agents and finds propagation and trust-capture failures
DefensePaperMicrosoft

Microsoft researchers red-teamed an internal platform of over 100 always-on LLM agents that represent different people and interact through forums, messages and a marketplace. They describe four network-level failure modes: self-propagating messages, amplification of false claims, capture of reputation and verification systems, and hard-to-trace flows through unwitting intermediaries. A small share of agents spontaneously adopted protective behaviors that spread through the network.

Dec 2025
Dec 9, 2025
OWASP publishes Top 10 for Agentic Applications (ASI01-ASI10)
PolicyStandardOWASP GenAI Security Project

The OWASP GenAI Security Project released its Top 10 for Agentic Applications, a list of ten risk categories specific to agents that plan, hold memory, call tools and act with delegated authority. The release came with an updated Agentic Threats and Mitigations taxonomy (v1.1) and a capture-the-flag practice platform.

Nov 2025
Nov 19, 2025
AppOmni shows second-order prompt injection recruiting privileged ServiceNow Now Assist agents
AttackVulnerability disclosureAppOmni, ServiceNow

AppOmni reports that instructions planted in an ordinary ServiceNow record could cause a low-privilege Now Assist agent to discover and task a more privileged agent, leading to record changes, data access and email exfiltration. The behavior follows default settings that group agents into teams and make them discoverable; ServiceNow called it intended and updated its documentation.

Aug 2025
Aug 14, 2025
NIST proposes SP 800-53 control overlays for securing AI, including single- and multi-agent systems
PolicyStandardNIST

NIST released a concept paper for Control Overlays for Securing AI Systems (COSAiS), which would tailor SP 800-53 security controls to AI use cases. The planned use cases include generative AI assistants, predictive AI, single-agent systems, multi-agent systems and controls for AI developers, informed by the AI 100-2 E2025 taxonomy. As of the project page, only an annotated outline for the predictive AI overlay (January 8, 2026) had followed; agent overlays had not been published.

May 2025
May 7, 2025
UC Santa Cruz study integrates LLM agents into CAGE 4 and finds RL defenders still outperform them
DefensePaperUC Santa Cruz

Researchers led by UC Santa Cruz integrated LLM agents into the CybORG CAGE 4 multi-agent defence environment and proposed a communication protocol for mixed LLM and RL teams. In their runs an all-RL team scored far better reward than an all-LLM (GPT-4o-mini) team and acted about 104 times faster, though the authors highlight LLM explainability and note the environment was designed for RL agents.

Feb 2025
Feb 17, 2025
OWASP Agentic Security Initiative releases Agentic AI Threats and Mitigations v1.0
PolicyFrameworkOWASP GenAI Security Project

OWASP's Agentic Security Initiative published a threat-model-based reference of emerging threats to LLM-powered autonomous agents and corresponding mitigations. It became the taxonomy underpinning the later OWASP Top 10 for Agentic Applications, which shipped with an updated v1.1 of this guide.

Feb 6, 2025
Cloud Security Alliance publishes MAESTRO seven-layer threat modeling framework for agentic AI
PolicyFrameworkCloud Security Alliance

The Cloud Security Alliance published MAESTRO (Multi-Agent Environment, Security, Threat, Risk, and Outcome), a threat modeling framework for agentic AI authored by Ken Huang. It organizes analysis into seven layers from foundation models to the agent ecosystem and highlights agent-specific threats such as goal manipulation, agent impersonation and collusion between agents.

Findings

Research openings

Benchmarks and tools