<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Capability thresholds · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/capability-thresholds/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/capability-thresholds/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on capability thresholds, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>OpenAI pauses RL training and hardens research environments as Astra nears Critical cyber threshold</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-pacing-development-cyber-critical-2026/</link>
<guid isPermaLink="false">event:openai-pacing-development-cyber-critical-2026</guid>
<pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>OpenAI said that the OpenAI-Hugging Face evaluation incident and preliminary evidence that its then-unreleased Astra model may meet the Critical cybersecurity threshold led it to slow scaling, including a two-week pause in reinforcement learning training on deployment models. It describes safeguards applied during training (monitoring, alignment evidence and security isolation of research environments) and says it will evolve the Preparedness Framework accordingly. It is a public case of a lab applying its Critical cyber threshold to development itself, including isolating its own training environments.</description>
</item>
<item>
<title>Executive Order 14409 creates classified cyber benchmarking for covered frontier models and a clearinghouse</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-eo-14409-frontier-ai-cyber-benchmarking-2026/</link>
<guid isPermaLink="false">event:us-eo-14409-frontier-ai-cyber-benchmarking-2026</guid>
<pubDate>Tue, 02 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Executive Order 14409 directs Treasury, NSA and CISA to develop a classified benchmarking process to assess advanced cyber capabilities of AI models and designate covered frontier models, with a voluntary framework for pre-release government and trusted-partner access. It also orders an AI cybersecurity clearinghouse to coordinate vulnerability scanning, validation and remediation with industry, and states it does not create mandatory licensing or pre-clearance. It is a US mechanism that designates models by cyber capability and gives the government early access before release to other trusted partners.</description>
</item>
<item>
<title>DeepMind Frontier Safety Framework v3.1 raises security for its cyber critical capability level</title>
<link>https://agentic-cyber-explorer.pages.dev/events/deepmind-frontier-safety-framework-v3-1-2026/</link>
<guid isPermaLink="false">event:deepmind-frontier-safety-framework-v3-1-2026</guid>
<pubDate>Fri, 17 Apr 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Google DeepMind's Frontier Safety Framework version 3.1 introduced Tracked Capability Levels for earlier warning in CBRN and ML R&amp;D and misalignment, and raised the recommended security for the CBRN, cyber and harmful manipulation CCLs to Security Level 2+. The cyber CCL definition itself, Cyber uplift level 1, is unchanged from v3.0. It raises the recommended weight security for cyber-capable models to Security Level 2+, adding protection against non-state and insider theft, while keeping the 2025 rationale that higher security levels are likely not warranted.</description>
</item>
<item>
<title>Anthropic RSP v3.0 rewrite adds risk reports and roadmaps; policy text does not name cyber</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-rsp-v3-rewrite-2026/</link>
<guid isPermaLink="false">event:anthropic-rsp-v3-rewrite-2026</guid>
<pubDate>Tue, 24 Feb 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic replaced its Responsible Scaling Policy with version 3.0, introducing Frontier Safety Roadmaps and Risk Reports and restating capability thresholds alongside recommended industry-wide mitigations. The published v3.0 policy document does not mention cyber capability; cyber safeguards for later models (Mythos, Fable 5) were described in separate announcements. Versions 3.1 through 3.4 followed between April and July 2026. Researchers tracking how labs gate cyber capability should note that Anthropic's cyber gating in 2026 ran through deployment decisions and classifiers rather than a written RSP cyber threshold.</description>
</item>
<item>
<title>Frontier Model Forum report sets out shared cyber thresholds for frontier AI safety frameworks</title>
<link>https://agentic-cyber-explorer.pages.dev/events/fmf-managing-advanced-cyber-risks-frameworks-2026/</link>
<guid isPermaLink="false">event:fmf-managing-advanced-cyber-risks-frameworks-2026</guid>
<pubDate>Fri, 13 Feb 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The Frontier Model Forum published a technical report on managing advanced cyber risks within frontier AI safety frameworks. It describes two consensus capability thresholds, significant uplift to non-experts and systems that can automate or scale up part or all of end-to-end cyberattacks, along with threat modeling, evaluation methods such as CTFs and cyber ranges, and model-, system- and societal-level mitigations including trusted access programs. It is the closest thing to an industry consensus definition of when a model's cyber capability should trigger stronger controls.</description>
</item>
<item>
<title>California SB 53 requires frontier AI frameworks covering autonomous cyberattack risk and incident reporting</title>
<link>https://agentic-cyber-explorer.pages.dev/events/california-sb53-frontier-ai-cyber-provisions-2025/</link>
<guid isPermaLink="false">event:california-sb53-frontier-ai-cyber-provisions-2025</guid>
<pubDate>Mon, 29 Sep 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>California's Transparency in Frontier AI Act (SB 53) requires large frontier developers to publish frontier AI frameworks addressing catastrophic risk, model weight cybersecurity and incident response, and to report critical safety incidents to the Office of Emergency Services. Its catastrophic risk definition includes a model engaging, with no meaningful human oversight, in conduct that is a cyberattack, where a single incident causes death or serious injury to more than 50 people or more than $1 billion in property damage. It is a binding US state law that ties catastrophic risk to autonomous cyberattack conduct by a model.</description>
</item>
<item>
<title>Google DeepMind Frontier Safety Framework v3 retains Cyber Uplift Level 1 critical capability level</title>
<link>https://agentic-cyber-explorer.pages.dev/events/deepmind-frontier-safety-framework-v3-cyber-ccl-2025/</link>
<guid isPermaLink="false">event:deepmind-frontier-safety-framework-v3-cyber-ccl-2025</guid>
<pubDate>Mon, 22 Sep 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Google DeepMind's Frontier Safety Framework version 3.0 adds a harmful manipulation CCL and expands misalignment and internal-deployment provisions. Its cyber domain keeps a single CCL, Cyber uplift level 1, for models providing sufficient uplift with high-impact cyber attacks to add expected harm at severe scale, paired with Security Level 2; the framework reasons that automated cyber-defense and social adaptation make higher security levels likely unwarranted. It shows a lab explicitly reasoning that cyber-capable model weights warrant only Security Level 2, the same as its CBRN and manipulation CCLs and below its ML R&amp;D CCLs, a choice revisited in 2026.</description>
</item>
<item>
<title>EU GPAI Code of Practice Safety and Security chapter lists cyber offence as a specified systemic risk</title>
<link>https://agentic-cyber-explorer.pages.dev/events/eu-gpai-code-of-practice-safety-security-2025/</link>
<guid isPermaLink="false">event:eu-gpai-code-of-practice-safety-security-2025</guid>
<pubDate>Thu, 10 Jul 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The European Commission received the final General-Purpose AI Code of Practice, whose Safety and Security chapter applies to providers of models with systemic risk under Article 55 of the AI Act. The chapter treats cyber offence as one of four specified systemic risks, requires a security goal covering non-state external and insider threats, and sets serious incident reporting deadlines that include five days for serious cybersecurity breaches. It is an operational EU text that commits signatory frontier providers to assess automated vulnerability discovery and exploit generation as a systemic risk.</description>
</item>
<item>
<title>Anthropic activates ASL-3 deployment and security protections for Claude Opus 4</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-asl3-activation-claude-opus-4-2025/</link>
<guid isPermaLink="false">event:anthropic-asl3-activation-claude-opus-4-2025</guid>
<pubDate>Thu, 22 May 2025 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>Anthropic activated ASL-3 protections for Claude Opus 4 as a precaution because it could not rule out ASL-3 CBRN risk; the announcement does not cite cyber capability as the trigger. The ASL-3 security standard it describes includes more than 100 controls to protect weights, two-party authorization for weight access, and egress bandwidth controls against exfiltration. It was a public activation of a higher safety level under a lab framework, and its egress and access controls are defenses against cyber theft of model weights.</description>
</item>
<item>
<title>OpenAI Preparedness Framework v2 sets High and Critical cybersecurity capability thresholds</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-preparedness-framework-v2-2025/</link>
<guid isPermaLink="false">event:openai-preparedness-framework-v2-2025</guid>
<pubDate>Tue, 15 Apr 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>OpenAI's Preparedness Framework version 2 makes cybersecurity one of three Tracked Categories and defines High and Critical capability thresholds, each tied to required safeguards. High covers automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities; Critical covers autonomous zero-day development across many hardened critical systems, and at Critical OpenAI commits to halt further development until adequate safeguards are specified. Its cyber thresholds are explicitly about autonomous, tool-augmented operation, so they are the operative gate for OpenAI's agentic cyber models in 2025-2026.</description>
</item>
<item>
<title>Anthropic RSP v2 lists cyber operations as a capability under ongoing assessment, not a threshold</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-rsp-v2-cyber-operations-assessment-2024/</link>
<guid isPermaLink="false">event:anthropic-rsp-v2-cyber-operations-assessment-2024</guid>
<pubDate>Tue, 15 Oct 2024 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic's Responsible Scaling Policy version 2.0, effective October 15, 2024, and its 2.x revisions list cyber operations among capabilities requiring ongoing assessment rather than as a formal capability threshold with required safeguards. The policy says Anthropic would consult cyber experts, consider tiered access controls or phased deployment for models with advanced cyber capabilities, and run pre- or post-deployment testing. It shows how one major lab treated autonomous exploit development as a monitored risk without committing to a cyber-specific ASL trigger during 2024-2025.</description>
</item>
<item>
<title>US Executive Order 14110 names offensive cyber capability as a dual-use foundation model risk</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-eo-14110-ai-cyber-provisions-2023/</link>
<guid isPermaLink="false">event:us-eo-14110-ai-cyber-provisions-2023</guid>
<pubDate>Mon, 30 Oct 2023 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>President Biden's executive order on safe, secure and trustworthy AI defined dual-use foundation models partly by their potential to enable offensive cyber operations through automated vulnerability discovery and exploitation. It required developers to report red-team results to the government and directed a federal pilot using AI to find and fix vulnerabilities in government systems. The order was revoked by Executive Order 14179 on January 23, 2025. It was a US executive instrument that treated automated vulnerability discovery and exploitation as a reportable frontier-model risk, and its revocation reset the US baseline.</description>
</item>
</channel>
</rss>
