Reference

Benchmarks and tools

The instruments of the field: what they measure, who maintains them, whether they are open, and which records cite them. Follow each link to the original for methods and results. Model families have their own profiles.

Benchmarks

NameWhat it coversMaintainerReleasedAccessRecords
Agent Red Teaming benchmarkRobustness of agents to adversarial prompts across target behaviors.Gray Swan AIJuly 2025restricted2
Agent Security BenchAttack success rates across injection, memory, and backdoor attack types.—October 2024open1
AgentDojoUtility and targeted attack success under injection, with and without defenses.ETH ZurichJune 2024open6
AutoPatchBenchPatch success on fuzzing crashes.MetaApril 2025open1
BountyBenchDetection, exploitation, and patching of real bug-bounty vulnerabilities.Stanford UniversityMay 2025open1
CTFusionContamination-resistant CTF performance.—May 2026unknown1
CTI-REALMDetection engineering from threat reports.MicrosoftMarch 2026open1
CVE-BenchAgent success against real, sandboxed web vulnerabilities.University of Illinois Urbana-ChampaignMarch 2025open0
CybenchAutonomous CTF solving across web, crypto, reverse engineering, forensics, and exploitation categories.Stanford UniversityAugust 2024open3
Cyber Defense BenchmarkShare of malicious events agents flag.Simbian AIApril 2026unknown1
CyberGymReproducing real vulnerabilities from descriptions and codebases.UC BerkeleyJune 2025open4
CyberSecEvalInsecure code generation, cyberattack helpfulness, prompt injection, and later defensive tasks.MetaDecember 2023open3
CyberSOCEvalSOC analysis tasks.Meta, CrowdStrikeSeptember 2025open1
ExCyTIn-BenchInvestigation accuracy for LLM agents.MicrosoftJuly 2025open1
ExploitBenchGraded exploitation progress, distinguishing crashes from code execution.—May 2026unknown2
ExploitGymWorking exploit generation.—May 2026unknown2
InjecAgentRate at which agents follow injected instructions.University of Illinois Urbana-ChampaignMarch 2024open1
InterCode-CTFInteractive CTF solving in a shell environment.—June 2023open0
MCPToxAttack success of poisoned MCP tool metadata.—August 2025open1
MonitoringBenchMonitor catch rates under refined attacks.—May 2026unknown1
NYU CTF BenchAutonomous CTF solving.New York UniversityJune 2024open0
PatchBenchPatch validity beyond crash reproduction.—September 2026open1
SandboxEscapeBenchContainer sandbox escape.UK AI Security InstituteMarch 2026unknown1
SEC-benchProof-of-concept generation and patching.—June 2025open1
SHADE-ArenaCovert sabotage success and monitor detection.AnthropicJune 2025open2

Tools and systems

NameWhat it coversMaintainerReleasedAccessRecords
Big SleepGoogle's LLM agent for finding vulnerabilities, developed from Project Naptime.Google Project Zero, Google DeepMindNovember 2024closed3
ButtercupTrail of Bits' open-sourced AIxCC cyber reasoning system.Trail of BitsAugust 2025open2
Claude Code SecurityAnthropic tool that scans code for vulnerabilities and proposes patches.AnthropicFebruary 2026restricted1
CodeMenderGoogle DeepMind agent that patches vulnerabilities and rewrites code to remove vulnerability classes.Google DeepMindOctober 2025closed1
Codex Security (formerly Aardvark)OpenAI agent that scans repositories, validates vulnerabilities, and proposes patches.OpenAIOctober 2025restricted2
Constitutional ClassifiersAutomated jailbreak success and over-refusal.AnthropicFebruary 2025closed2
LlamaFirewallAttack success reduction and utility cost on AgentDojo.MetaMay 2025open1
mcp-scanScanner that checks installed MCP servers for tool poisoning and related risks.Invariant LabsApril 2025open1
Microsoft Security CopilotMicrosoft's generative AI assistant and agents for security operations.MicrosoftMarch 2023closed3
OSS-CRSPackaging that makes AIxCC cyber reasoning systems runnable locally.OpenSSFMarch 2026open2
PageBreakGoogle agent that finds web vulnerabilities in its own applications with deterministic validation.GoogleSeptember 2026closed1
Project IreMicrosoft agent that reverse engineers and classifies software as malicious or benign.MicrosoftAugust 2025closed1

Datasets

NameWhat it coversMaintainerReleasedAccessRecords
ARVODataset of reproducible OSS-Fuzz vulnerabilities with located fixes.Arizona State UniversityAugust 2024open1
LLMail-InjectAdaptive attacker success against layered defenses.MicrosoftJune 2025open1

Environments

NameWhat it coversMaintainerReleasedAccessRecords
CAGE Challenge 4 / CybORGDefender performance and service continuity.The Technical Cooperation ProgramFebruary 2024open2
ControlArenaSafety and usefulness of control protocols.UK AI Security InstituteOctober 2025open1

Frameworks

NameWhat it coversMaintainerReleasedAccessRecords
CaMeLProvable protection against control-flow hijacking; utility cost on AgentDojo.Google DeepMindMarch 2025open1
Cyber Jailbreak Severity frameworkSeverity of cyber safeguard bypasses.AnthropicJuly 2026open1
InspectUK AISI's open-source framework for running AI evaluations, used by many cyber evaluations.UK AI Security InstituteMay 2024open3
Instruction hierarchyRobustness to injection and jailbreaks.OpenAIApril 2024closed2
Model Context ProtocolOpen protocol for connecting AI applications to tools and data sources.Model Context Protocol projectNovember 2024open6
SecAlignInjection success and utility.UC Berkeley, MetaOctober 2024open2
SpotlightingInjection success under static and adaptive attacks.MicrosoftMarch 2024open4
StruQInjection success against static attacks.UC BerkeleyFebruary 2024open2

Programs

NameWhat it coversMaintainerReleasedAccessRecords
DARPA AI Cyber Challenge (AIxCC)Two-year DARPA competition for AI systems that find and fix vulnerabilities in open-source software.DARPA, ARPA-HAugust 2023open5
DARPA CASTLEDARPA program training reinforcement-learning agents for autonomous network defense.DARPAJuly 2024closed1
DARPA DICEDARPA program on decentralized agent collectives that stay controllable when members are lost or compromised.DARPAJune 2026closed1
OSS-FuzzGoogle's continuous fuzzing service for open-source software, which added LLM-generated fuzz targets in 2023.Google Open Source Security TeamDecember 2016open4
Project GlasswingAnthropic program giving vetted defenders early access to a cyber-capable model to find and fix vulnerabilities.AnthropicApril 2026restricted1