Locate

Attack surface atlas

A reference architecture of an agentic system. Each component collects the documented attacks and incidents against it, the defense research that addresses it, and the findings about both. Gaps between the two are research openings.

attacks and incidentsdefense researchfindingsbox tint = share of documented attacks
Supply chainweights · packages · extensions1323Access gatewho may use the model041Untrusted contentweb · email · files · issues28278Agent modelplanner · policy2235Memory & contextpersistent memory · compaction511GAPTools & MCPAPIs · connectors · MCP282910Credentialstokens · secrets · scopes1973GAPExecution sandboxcode · network egress16105Other agentsdelegation · shared state553Monitorautomated oversight1133Human approverescalation · approval8206GAPEvaluation environmenttest ranges · harnesses8334
web · email · files · issues
Untrusted content

Anything the agent reads that an outsider could write. The entry point for indirect prompt injection.

Everything on untrusted content →

Components

ComponentWhat it coversAttacksDefense recordsMeasured defense findings
Supply chainModel artifacts, packages, IDE extensions, and skills that an agent system is built from.1321
Access gateAccount controls, trusted access programs, and verification that decide who can use cyber capability.040
Untrusted contentAnything the agent reads that an outsider could write. The entry point for indirect prompt injection.28274
Agent modelThe model making decisions: its safeguards, refusals, and susceptibility to jailbreaks.2232
Memory & contextState the agent carries between steps and sessions, including summaries it writes for itself.510
Tools & MCPThe actions an agent can take and the servers that provide them.28293
CredentialsThe authority an agent holds, and what it can reach with it.1970
Execution sandboxWhere agent-written code runs and what that environment can reach.16101
Other agentsSub-agents, peer agents, and the messages passed between them.552
MonitorAutomated systems that watch agent behavior and can stop it.1132
Human approverPeople asked to approve, review, or intervene, and the load placed on them.8200
Evaluation environmentThe environments used to test agents, which have themselves become part of the attack surface.8332