Topics/Offense & defense

Autonomous pentesting

Agents that run offensive security work end to end.

2 records0 findings0 openings0 benchmarks and toolsLatest record
RangeLanes
2 of 2 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
2026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

2 records · newest first
Aug 2026
Aug 31, 2026
MITRE ATLAS adds autonomous attack techniques and case studies of agent-driven intrusions
PolicyStandardMITRE

MITRE's August 2026 ATLAS release added techniques describing AI agents acting as attackers, including autonomous reconnaissance, attack-path adaptation, attack orchestration and autonomous exploit development. It also added agent-control mitigations and case studies including the GTG-1002 Claude Code espionage campaign and autonomous OpenAI evaluation agents compromising Hugging Face infrastructure.

Apr 2025
Apr 15, 2025
OpenAI Preparedness Framework v2 sets High and Critical cybersecurity capability thresholds
PolicyFrameworkOpenAI

OpenAI's Preparedness Framework version 2 makes cybersecurity one of three Tracked Categories and defines High and Critical capability thresholds, each tied to required safeguards. High covers automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities; Critical covers autonomous zero-day development across many hardened critical systems, and at Critical OpenAI commits to halt further development until adequate safeguards are specified.