Reference
Sources
344 citations from 135 publishers. 87% are primary sources.
arXiv 56
- Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
- It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center
- PatchBench: Evaluating AI Agents for Vulnerability Patching
- SoK: DARPA's AI Cyber Challenge (AIxCC) (HTML, v5)
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- Measuring Security Without Fooling Ourselves (HTML)
- Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- MonitoringBench (HTML)
- MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
- Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
- How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
- CTI-REALM: Benchmark to Evaluate Agent Performance on Security Detection Rule Generation Capabilities
- OSS-CRS: Liberating AIxCC Cyber Reasoning Systems for Real-World Open-Source Security
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape (HTML)
- Quantifying Frontier LLM Capabilities for Container Sandbox Escape
- SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned
- Randomized Controlled Trials for Phishing Triage Agent
- The Attacker Moves Second (HTML)
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions (v3)
- CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
- EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- Large Language Models are Autonomous Cyber Defenders (v2 HTML)
- ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems (v2)
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
- LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
- Design Patterns for Securing LLM Agents against Prompt Injections (HTML)
- Design Patterns for Securing LLM Agents against Prompt Injections
- BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
- Lessons from Defending Gemini Against Indirect Prompt Injections (HTML)
- Lessons from Defending Gemini Against Indirect Prompt Injections
- Large Language Models are Autonomous Cyber Defenders
- LlamaFirewall (HTML)
- LlamaFirewall: An open source guardrail system for building secure AI agents
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
- Defeating Prompt Injections by Design
- Multi-Objective Reinforcement Learning for Automated Resilient Cyber Defence (preprint)
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
- AgentDojo (HTML, current version)
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
- InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
- StruQ: Defending Against Prompt Injection with Structured Queries
- AI Control: Improving Safety Despite Intentional Subversion
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Anthropic 24
- Investigating three real-world incidents in our cybersecurity evaluations
- More details on Fable 5's cyber safeguards and our jailbreak framework
- Redeploying Fable 5
- Statement on the US government directive to suspend access to Fable 5 and Mythos 5
- What we learned mapping a year’s worth of AI-enabled cyber threats
- Project Glasswing: An initial update
- Project Glasswing: Securing critical software for the AI era
- Responsible Scaling Policy Version 3.0
- Making frontier cybersecurity capabilities available to defenders
- System Card: Claude Opus 4.6
- Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks
- Experimenting with AI to defend critical infrastructure
- Mitigating the risk of prompt injections in browser use
- Disrupting the first reported AI-orchestrated cyber espionage campaign (full report)
- Disrupting the first reported AI-orchestrated cyber espionage campaign
- Beyond permission prompts: making Claude Code more secure and autonomous
- Building AI for cyber defenders
- Detecting and countering misuse of AI: August 2025
- Piloting Claude in Chrome
- Activating AI Safety Level 3 Protections
- Anthropic's Responsible Scaling Policy (version 2.2)
- Constitutional Classifiers: Defending against universal jailbreaks
- Anthropic's Responsible Scaling Policy (updates)
- Anthropic's Responsible Scaling Policy (updates)
OpenAI 18
- Self-generated prompt injections in compaction summaries
- Signing up for disposable emails and searching GitHub for leaked API keys
- Misalignment Reports and Notices
- The Hugging Face incident and the road ahead
- Pacing model development in an era of cyber-critical capabilities
- Third-party cyber evaluations involving OpenAI models
- OpenAI and Hugging Face partner to address security incident during model evaluation
- A shared playbook for trustworthy third party evaluations
- How we monitor internal coding agents for misalignment
- Designing AI agents to resist prompt injection
- Improving instruction hierarchy in frontier LLMs
- Codex Security: now in research preview
- Introducing Lockdown Mode and Elevated Risk labels in ChatGPT
- Keeping your data safe when an AI agent clicks a link
- Continuously hardening ChatGPT Atlas against prompt injection attacks
- Understanding prompt injections: a frontier security challenge
- Introducing Aardvark: OpenAI's agentic security researcher
- Preparedness Framework, Version 2
NIST 14
- NIST Mathematical Proof Supports Transition to a Continuous-Monitor-and-Update Security Model for AI Systems
- CAISI Evaluation of DeepSeek V4 Pro
- Insights into AI Agent Security from a Large-Scale Red-Teaming Competition
- Announcing the "AI Agent Standards Initiative" for Interoperable and Secure Innovation
- Draft NIST Guidelines Rethink Cybersecurity for the AI Era
- CAISI Evaluation of DeepSeek AI Models Finds Shortcomings and Risks
- NIST Releases Control Overlays for Securing AI Systems Concept Paper
- NIST AI 100-2 E2025 (PDF)
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2 E2025)
- Technical Blog: Strengthening AI Agent Hijacking Evaluations
- Updated Guidelines for Managing Misuse Risk for Dual-Use Foundation Models
- NIST AI 100-2 E2023 (PDF)
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2 E2023)
- Center for AI Standards and Innovation (CAISI)
UK National Cyber Security Centre 11
- The AI shift in cyber risk: why leaders must act now
- Thinking carefully before adopting agentic AI
- 10 questions to ask when using AI models to find vulnerabilities
- Retaining defensive advantage in the age of frontier AI cyber capabilities
- Why cyber defenders need to be ready for frontier AI
- Prompt injection is not SQL injection (it may be worse)
- Impact of AI on cyber threat from now to 2027
- The near-term impact of AI on the cyber threat
- Guidelines for secure AI system development
- Frontier AI: what you need to know
- Frontier AI: what you need to know
The Hacker News 8
- Attacker Hijacks AI Coding Assistant Session, Spreads Shai-Hulud Across About 100 Repositories
- OpenAI Agents Linked to RubyGems Campaign That Gained RCE on RubyDoc Servers
- Microsoft Copilot Personal Flaws Could Let One Click Exfiltrate Data From Connected Apps
- Critical Cursor Flaws Could Let Prompt Injection Escape Sandbox and Run Commands
- Anthropic MCP Design Vulnerability Enables RCE, Threatening AI Supply Chain
- Actively Exploited nginx-ui Flaw (CVE-2026-33032) Enables Full Nginx Server Takeover
- Three Flaws in Anthropic MCP Git Server Enable File Access and Code Execution
- First Malicious MCP Server Found Stealing Emails in Rogue Postmark-MCP Package
UK AI Security Institute 7
- Incident Report: unsanctioned agent behaviour during cyber testing
- UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
- How our Control Red Team is stress-testing frontier monitors
- Cheating behaviour in frontier model evaluations
- More compute, more capability: Why AI agent evaluations need to account for test-time compute
- How fast is autonomous AI cyber capability advancing?
- Introducing ControlArena: A library for running AI control experiments
DARPA 6
- DICE: Decentralized Artificial Intelligence through Controlled Emergence
- AI Cyber Challenge marks pivotal inflection point for cyber defense
- DARPA AI Cyber Challenge Proves Promise of AI-Driven Cybersecurity
- DARPA to Bring AI Cyber Challenge Semifinal Competition to DEF CON 32
- DARPA AI Cyber Challenge Aims to Secure Nation's Most Critical Software
- CASTLE: Cyber Agents for Security Testing and Learning Environments
The White House 6
- Fact Sheet: President Donald J. Trump Promotes Advanced Artificial Intelligence Innovation and Security
- Promoting Advanced Artificial Intelligence Innovation and Security
- President Trump’s Cyber Strategy for America
- White House Unveils America's AI Action Plan
- Winning the Race: America's AI Action Plan
- Removing Barriers to American Leadership in Artificial Intelligence
CISA 5
- Careful Adoption of Agentic AI Services
- CISA, US and International Partners Release Guide to Secure Adoption of Agentic AI
- Principles for the Secure Integration of Artificial Intelligence in Operational Technology
- New Best Practices Guide for Securing AI Data Released
- Joint Guidance on Deploying AI Systems Securely
Google DeepMind 5
OWASP GenAI Security Project 5
Simon Willison's Weblog 5
- OpenAI agents attacked RubyGems back in May
- OpenAI's accidental cyberattack against Hugging Face is science fiction that happened
- Google Antigravity Exfiltrates Data
- New prompt injection papers: Agents Rule of Two and The Attacker Moves Second
- The lethal trifecta for AI agents: private data, untrusted content, and external communication
European Commission 4
Fortune 4
- OpenAI's AI agents secretly used a German wiki website as a message board. OpenAI stayed quiet about it for weeks.
- OpenAI says its AI models secretly broke out of a secure test environment and hacked into AI company Hugging Face in order to cheat on an evaluation
- Anthropic disables Fable and Mythos AI models following U.S. government export ban
- An AI-powered coding tool wiped out a software company's database in 'catastrophic failure'