Topics/Governance

Capability thresholds

Lab and government thresholds that trigger safeguards or gate release.

12 records0 findings0 openings0 benchmarks and toolsLatest record
RangeLanes
10 of 12 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
202420252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

10 records · newest first
Aug 2026
Aug 18, 2026
OpenAI pauses RL training and hardens research environments as Astra nears Critical cyber threshold
PolicyFrameworkOpenAI

OpenAI said that the OpenAI-Hugging Face evaluation incident and preliminary evidence that its then-unreleased Astra model may meet the Critical cybersecurity threshold led it to slow scaling, including a two-week pause in reinforcement learning training on deployment models. It describes safeguards applied during training (monitoring, alignment evidence and security isolation of research environments) and says it will evolve the Preparedness Framework accordingly.

Jun 2026
Jun 2, 2026
Executive Order 14409 creates classified cyber benchmarking for covered frontier models and a clearinghouse
PolicyRegulationThe White House, US Department of the Treasury, NSA

Executive Order 14409 directs Treasury, NSA and CISA to develop a classified benchmarking process to assess advanced cyber capabilities of AI models and designate covered frontier models, with a voluntary framework for pre-release government and trusted-partner access. It also orders an AI cybersecurity clearinghouse to coordinate vulnerability scanning, validation and remediation with industry, and states it does not create mandatory licensing or pre-clearance.

Apr 2026
Feb 2026
Feb 24, 2026
Anthropic RSP v3.0 rewrite adds risk reports and roadmaps; policy text does not name cyber
PolicyFrameworkAnthropic

Anthropic replaced its Responsible Scaling Policy with version 3.0, introducing Frontier Safety Roadmaps and Risk Reports and restating capability thresholds alongside recommended industry-wide mitigations. The published v3.0 policy document does not mention cyber capability; cyber safeguards for later models (Mythos, Fable 5) were described in separate announcements. Versions 3.1 through 3.4 followed between April and July 2026.

Feb 13, 2026
Frontier Model Forum report sets out shared cyber thresholds for frontier AI safety frameworks
PolicyFrameworkFrontier Model Forum

The Frontier Model Forum published a technical report on managing advanced cyber risks within frontier AI safety frameworks. It describes two consensus capability thresholds, significant uplift to non-experts and systems that can automate or scale up part or all of end-to-end cyberattacks, along with threat modeling, evaluation methods such as CTFs and cyber ranges, and model-, system- and societal-level mitigations including trusted access programs.

Sep 2025
Sep 29, 2025
California SB 53 requires frontier AI frameworks covering autonomous cyberattack risk and incident reporting
PolicyRegulationState of California

California's Transparency in Frontier AI Act (SB 53) requires large frontier developers to publish frontier AI frameworks addressing catastrophic risk, model weight cybersecurity and incident response, and to report critical safety incidents to the Office of Emergency Services. Its catastrophic risk definition includes a model engaging, with no meaningful human oversight, in conduct that is a cyberattack, where a single incident causes death or serious injury to more than 50 people or more than $1 billion in property damage.

Sep 22, 2025
Google DeepMind Frontier Safety Framework v3 retains Cyber Uplift Level 1 critical capability level
PolicyFrameworkGoogle DeepMind

Google DeepMind's Frontier Safety Framework version 3.0 adds a harmful manipulation CCL and expands misalignment and internal-deployment provisions. Its cyber domain keeps a single CCL, Cyber uplift level 1, for models providing sufficient uplift with high-impact cyber attacks to add expected harm at severe scale, paired with Security Level 2; the framework reasons that automated cyber-defense and social adaptation make higher security levels likely unwarranted.

Jul 2025
Jul 10, 2025
EU GPAI Code of Practice Safety and Security chapter lists cyber offence as a specified systemic risk
PolicyFrameworkEuropean Commission

The European Commission received the final General-Purpose AI Code of Practice, whose Safety and Security chapter applies to providers of models with systemic risk under Article 55 of the AI Act. The chapter treats cyber offence as one of four specified systemic risks, requires a security goal covering non-state external and insider threats, and sets serious incident reporting deadlines that include five days for serious cybersecurity breaches.

May 2025
May 22, 2025
Anthropic activates ASL-3 deployment and security protections for Claude Opus 4
CapabilityThresholdAnthropic

Anthropic activated ASL-3 protections for Claude Opus 4 as a precaution because it could not rule out ASL-3 CBRN risk; the announcement does not cite cyber capability as the trigger. The ASL-3 security standard it describes includes more than 100 controls to protect weights, two-party authorization for weight access, and egress bandwidth controls against exfiltration.

Apr 2025
Apr 15, 2025
OpenAI Preparedness Framework v2 sets High and Critical cybersecurity capability thresholds
PolicyFrameworkOpenAI

OpenAI's Preparedness Framework version 2 makes cybersecurity one of three Tracked Categories and defines High and Critical capability thresholds, each tied to required safeguards. High covers automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities; Critical covers autonomous zero-day development across many hardened critical systems, and at Critical OpenAI commits to halt further development until adequate safeguards are specified.