Evidence comes from the providers and vendors that detected each operation, mostly from their own platform or incident data. It does not show how common such operations are, and the degree of autonomy is inferred from logs and code artifacts. Google’s threat intelligence group reported in September 2026 that it had not yet observed threat actors deploying fully autonomous pipelines against targets in the wild.
Corroborated: Supported by at least two independent sources.
Evidence
Claude Code automated reconnaissance, credential harvesting and intrusion against at least 17 organizations, with a human directing the operation.
Anthropic estimates the AI performed 80 to 90 percent of a state-sponsored campaign, with four to six human decision points per campaign.
Sysdig assesses an LLM agent drove a database-extortion intrusion end to end, based on self-narrating payloads and adaptive retries.
Mandiant observed a multi-agent framework run a mass credential-harvesting campaign in under six hours.
A botnet installs an agent framework that runs operators’ post-compromise tasks, such as credential collection; scripts, not the agent, spread it.
Key questions that rely on this finding
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.