Reference
Benchmarks and tools
The instruments of the field: what they measure, who maintains them, whether they are open, and which records cite them. Follow each link to the original for methods and results. Model families have their own profiles.
Benchmarks
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| Agent Red Teaming benchmark | Robustness of agents to adversarial prompts across target behaviors. | Gray Swan AI | July 2025 | restricted | 2 |
| Agent Security Bench | Attack success rates across injection, memory, and backdoor attack types. | — | October 2024 | open | 1 |
| AgentDojo | Utility and targeted attack success under injection, with and without defenses. | ETH Zurich | June 2024 | open | 6 |
| AutoPatchBench | Patch success on fuzzing crashes. | Meta | April 2025 | open | 1 |
| BountyBench | Detection, exploitation, and patching of real bug-bounty vulnerabilities. | Stanford University | May 2025 | open | 1 |
| CTFusion | Contamination-resistant CTF performance. | — | May 2026 | unknown | 1 |
| CTI-REALM | Detection engineering from threat reports. | Microsoft | March 2026 | open | 1 |
| CVE-Bench | Agent success against real, sandboxed web vulnerabilities. | University of Illinois Urbana-Champaign | March 2025 | open | 0 |
| Cybench | Autonomous CTF solving across web, crypto, reverse engineering, forensics, and exploitation categories. | Stanford University | August 2024 | open | 3 |
| Cyber Defense Benchmark | Share of malicious events agents flag. | Simbian AI | April 2026 | unknown | 1 |
| CyberGym | Reproducing real vulnerabilities from descriptions and codebases. | UC Berkeley | June 2025 | open | 4 |
| CyberSecEval | Insecure code generation, cyberattack helpfulness, prompt injection, and later defensive tasks. | Meta | December 2023 | open | 3 |
| CyberSOCEval | SOC analysis tasks. | Meta, CrowdStrike | September 2025 | open | 1 |
| ExCyTIn-Bench | Investigation accuracy for LLM agents. | Microsoft | July 2025 | open | 1 |
| ExploitBench | Graded exploitation progress, distinguishing crashes from code execution. | — | May 2026 | unknown | 2 |
| ExploitGym | Working exploit generation. | — | May 2026 | unknown | 2 |
| InjecAgent | Rate at which agents follow injected instructions. | University of Illinois Urbana-Champaign | March 2024 | open | 1 |
| InterCode-CTF | Interactive CTF solving in a shell environment. | — | June 2023 | open | 0 |
| MCPTox | Attack success of poisoned MCP tool metadata. | — | August 2025 | open | 1 |
| MonitoringBench | Monitor catch rates under refined attacks. | — | May 2026 | unknown | 1 |
| NYU CTF Bench | Autonomous CTF solving. | New York University | June 2024 | open | 0 |
| PatchBench | Patch validity beyond crash reproduction. | — | September 2026 | open | 1 |
| SandboxEscapeBench | Container sandbox escape. | UK AI Security Institute | March 2026 | unknown | 1 |
| SEC-bench | Proof-of-concept generation and patching. | — | June 2025 | open | 1 |
| SHADE-Arena | Covert sabotage success and monitor detection. | Anthropic | June 2025 | open | 2 |
Tools and systems
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| Big Sleep | Google's LLM agent for finding vulnerabilities, developed from Project Naptime. | Google Project Zero, Google DeepMind | November 2024 | closed | 3 |
| Buttercup | Trail of Bits' open-sourced AIxCC cyber reasoning system. | Trail of Bits | August 2025 | open | 2 |
| Claude Code Security | Anthropic tool that scans code for vulnerabilities and proposes patches. | Anthropic | February 2026 | restricted | 1 |
| CodeMender | Google DeepMind agent that patches vulnerabilities and rewrites code to remove vulnerability classes. | Google DeepMind | October 2025 | closed | 1 |
| Codex Security (formerly Aardvark) | OpenAI agent that scans repositories, validates vulnerabilities, and proposes patches. | OpenAI | October 2025 | restricted | 2 |
| Constitutional Classifiers | Automated jailbreak success and over-refusal. | Anthropic | February 2025 | closed | 2 |
| LlamaFirewall | Attack success reduction and utility cost on AgentDojo. | Meta | May 2025 | open | 1 |
| mcp-scan | Scanner that checks installed MCP servers for tool poisoning and related risks. | Invariant Labs | April 2025 | open | 1 |
| Microsoft Security Copilot | Microsoft's generative AI assistant and agents for security operations. | Microsoft | March 2023 | closed | 3 |
| OSS-CRS | Packaging that makes AIxCC cyber reasoning systems runnable locally. | OpenSSF | March 2026 | open | 2 |
| PageBreak | Google agent that finds web vulnerabilities in its own applications with deterministic validation. | September 2026 | closed | 1 | |
| Project Ire | Microsoft agent that reverse engineers and classifies software as malicious or benign. | Microsoft | August 2025 | closed | 1 |
Datasets
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| ARVO | Dataset of reproducible OSS-Fuzz vulnerabilities with located fixes. | Arizona State University | August 2024 | open | 1 |
| LLMail-Inject | Adaptive attacker success against layered defenses. | Microsoft | June 2025 | open | 1 |
Environments
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| CAGE Challenge 4 / CybORG | Defender performance and service continuity. | The Technical Cooperation Program | February 2024 | open | 2 |
| ControlArena | Safety and usefulness of control protocols. | UK AI Security Institute | October 2025 | open | 1 |
Frameworks
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| CaMeL | Provable protection against control-flow hijacking; utility cost on AgentDojo. | Google DeepMind | March 2025 | open | 1 |
| Cyber Jailbreak Severity framework | Severity of cyber safeguard bypasses. | Anthropic | July 2026 | open | 1 |
| Inspect | UK AISI's open-source framework for running AI evaluations, used by many cyber evaluations. | UK AI Security Institute | May 2024 | open | 3 |
| Instruction hierarchy | Robustness to injection and jailbreaks. | OpenAI | April 2024 | closed | 2 |
| Model Context Protocol | Open protocol for connecting AI applications to tools and data sources. | Model Context Protocol project | November 2024 | open | 6 |
| SecAlign | Injection success and utility. | UC Berkeley, Meta | October 2024 | open | 2 |
| Spotlighting | Injection success under static and adaptive attacks. | Microsoft | March 2024 | open | 4 |
| StruQ | Injection success against static attacks. | UC Berkeley | February 2024 | open | 2 |
Programs
| Name | What it covers | Maintainer | Released | Access | Records |
|---|---|---|---|---|---|
| DARPA AI Cyber Challenge (AIxCC) | Two-year DARPA competition for AI systems that find and fix vulnerabilities in open-source software. | DARPA, ARPA-H | August 2023 | open | 5 |
| DARPA CASTLE | DARPA program training reinforcement-learning agents for autonomous network defense. | DARPA | July 2024 | closed | 1 |
| DARPA DICE | DARPA program on decentralized agent collectives that stay controllable when members are lost or compromised. | DARPA | June 2026 | closed | 1 |
| OSS-Fuzz | Google's continuous fuzzing service for open-source software, which added LLM-generated fuzz targets in 2023. | Google Open Source Security Team | December 2016 | open | 4 |
| Project Glasswing | Anthropic program giving vetted defenders early access to a cyber-capable model to find and fix vulnerabilities. | Anthropic | April 2026 | restricted | 1 |