How it works
An operator gives an agent tools such as scanners, shells and cloud APIs, plus task instructions that are often split into small steps that look harmless on their own, so safeguards do not see the whole operation. The agent runs in a loop: it issues commands, reads the results, decides the next step and reports back, which lets one operator work many targets in parallel at machine speed.
Reported cases come mostly from the provider or vendor that detected them, so prevalence is unknown. Anthropic reports its espionage case succeeded against only a small number of about thirty targets and that the agent hallucinated results; Google reported in 2026 that it had not yet seen fully autonomous pipelines in the wild. Attributing autonomy often rests on artifacts such as self-narrating code or the timing of operations.
What we know
1 corroboratedRecords over time
Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.
Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.
Records in view
5 records · newest firstGoogle Threat Intelligence Group's September 2026 tracker, drawing on Mandiant incident response, reports adversaries shifting from basic prompting to agentic workflows. In one case a suspected financially motivated actor used an AI coding chatbot and agent instruction files on compromised cloud infrastructure to build and run a mass credential-harvesting campaign in under six hours, compromising thousands of third-party credentials. GTIG also reports attackers targeting AI coding assistants and LLM security scanners in software supply-chain compromises, theft of proprietary AI models and data, and a growing underground market for AI accounts.
Sysdig's threat research team reports an operator it calls JADEPUFFER that gained access through a vulnerability in an internet-facing Langflow server (CVE-2025-3248), harvested credentials on that host, then used root database credentials of unknown origin against a separate production database server and ran a database-extortion playbook. Sysdig assesses the operation was driven end to end by an LLM agent, citing self-narrating payloads with natural-language reasoning and rapid adaptive retries, and calls it the first documented case of agentic ransomware.
Anthropic analyzed 832 accounts it banned for malicious cyber activity between March 2025 and March 2026 and mapped their use of Claude onto MITRE ATT&CK. It reports that the most common AI use was preparation such as writing malware, that use shifted toward activity after initial compromise, and that the share of actors its system rated medium risk or higher rose from 33% to 56% between the two six-month halves.
Anthropic reports that in mid-September 2025 a group it assesses with high confidence to be Chinese state-sponsored used Claude Code inside an attack framework to attempt intrusions into about thirty organizations, succeeding in a small number. The operators got past safeguards by splitting the work into innocuous-looking tasks and claiming to be a security firm doing defensive testing; Anthropic says the AI performed 80 to 90 percent of the campaign, with people at a handful of decision points.
Anthropic's August 2025 threat intelligence report describes a criminal who used Claude Code to automate reconnaissance, credential harvesting and network intrusion against at least 17 organizations, including healthcare, emergency services, government and religious institutions, then threatened to publish the stolen data. The report also describes North Korean operatives using Claude to obtain and keep remote technical jobs, and a low-skill actor selling ransomware developed with Claude.