Topics/Governance

Access controls

Trusted access programs, tiered access, and verification for cyber use.

12 records0 findings0 openings1 benchmarks and toolsLatest record
RangeLanes
11 of 12 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
20252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

11 records · newest first
Sep 2026
Sep 2, 2026
Google releases Gemini 3.8 Flash Cyber for trusted defenders, emphasizing automated patching
DefenseTool releaseGoogle, Google DeepMind, Collinear

Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-tuned model with more permissive cyber mitigations, available only to trusted defenders through a new Fairwind Program. Google says it prioritized vulnerability fixing over exploitation and reports 47.2% pass@1 on Collinear's CWE-Bench patching benchmark, over 70% on an internal 20-language discovery benchmark, and 2.6 times more correct Chrome patches than larger commercial models.

Jul 2026
Jul 7, 2026
European Commission presents EU Action Plan on Cybersecurity and Artificial Intelligence
PolicyProgramEuropean Commission, ENISA

The Commission presented an action plan responding to advanced AI models that can both improve and undermine cybersecurity. It plans an EU capacity to evaluate AI models, a European blueprint for structured access to advanced AI capabilities developed with ENISA, a secure ENISA-JRC platform to test AI for cybersecurity, AI-assisted vulnerability fixing, and a campaign to secure critical open-source software. ENISA published its own recommendations for the frontier AI era the same day.

Jul 2, 2026
Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners
PolicyFrameworkAnthropic

Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response.

Jun 2026
Jun 30, 2026
US lifts export controls on Fable 5 and Mythos 5; Anthropic redeploys with new cyber classifier
PolicyRegulationAnthropic, US Department of Commerce, US Center for AI Standards and Innovation

Anthropic announced that export controls on Fable 5 and Mythos 5 had been lifted and that Fable 5 would be redeployed globally from July 1, 2026 with an improved safety classifier. Anthropic says the classifier blocks the technique described in an Amazon report in over 99% of cases and that CAISI researchers tested its prior and new safeguards. Mythos 5 access was restored for a set of US organizations after government approval on June 26.

Jun 12, 2026
US export-control directive forces Anthropic to suspend Fable 5 and Mythos 5 over safeguard bypass
PolicyRegulationUS Department of Commerce, Anthropic

Anthropic said the US government issued an export control directive, citing national security authorities, barring access to Fable 5 and Mythos 5 by foreign nationals, after officials said they had found a way to jailbreak Fable 5's safeguards. Anthropic said the net effect was that it had to disable both models for all customers to comply, while other Claude models stayed available. Anthropic disputed the rationale, arguing the demonstrated vulnerabilities were minor and that the standard applied industry-wide would halt new frontier deployments.

Jun 2, 2026
Executive Order 14409 creates classified cyber benchmarking for covered frontier models and a clearinghouse
PolicyRegulationThe White House, US Department of the Treasury, NSA

Executive Order 14409 directs Treasury, NSA and CISA to develop a classified benchmarking process to assess advanced cyber capabilities of AI models and designate covered frontier models, with a voluntary framework for pre-release government and trusted-partner access. It also orders an AI cybersecurity clearinghouse to coordinate vulnerability scanning, validation and remediation with industry, and states it does not create mandatory licensing or pre-clearance.

May 2026
May 22, 2026
Glasswing update: over 10,000 high-severity bugs found, but only 75 of 530 disclosed OSS bugs patched
DefenseProgramAnthropic, Cloudflare, Mozilla

Anthropic reports that about 50 Glasswing partners used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities, and that its own scan of over 1,000 open-source projects produced 6,202 model-estimated high/critical findings. Of 1,752 assessed, mostly by six independent firms, 90.6% were true positives; Anthropic estimates 530 high/critical bugs disclosed, of which 75 were patched, and says triage and patching capacity, not discovery, is the bottleneck.

Apr 2026
Apr 7, 2026
Anthropic launches Project Glasswing to give defenders early access to Claude Mythos Preview
DefenseProgramAnthropic, Amazon Web Services, Apple

Anthropic launched Project Glasswing with AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks to use the unreleased Claude Mythos Preview for defensive security work, extending access to over 40 more organizations that maintain critical software. Anthropic committed up to $100M in usage credits and $4M in donations to open-source security groups, and reports Mythos Preview found thousands of high-severity vulnerabilities, including in every major operating system and browser.

Feb 2026
Feb 13, 2026
OpenAI adds Lockdown Mode and Elevated Risk labels to ChatGPT to limit prompt injection exfiltration
DefenseTool releaseOpenAI

OpenAI introduced Lockdown Mode, an optional setting that deterministically disables or limits capabilities an attacker could exploit through prompt injection, such as live web access, image support in responses, Deep Research, Agent Mode, live connectors and file downloads. Elevated Risk labels flag network-related features in ChatGPT, Atlas and Codex that carry extra risk. Lockdown Mode first launched for enterprise-type plans, and a June 4, 2026 update says it is rolling out to personal and self-serve Business accounts.

Feb 13, 2026
Frontier Model Forum report sets out shared cyber thresholds for frontier AI safety frameworks
PolicyFrameworkFrontier Model Forum

The Frontier Model Forum published a technical report on managing advanced cyber risks within frontier AI safety frameworks. It describes two consensus capability thresholds, significant uplift to non-experts and systems that can automate or scale up part or all of end-to-end cyberattacks, along with threat modeling, evaluation methods such as CTFs and cyber ranges, and model-, system- and societal-level mitigations including trusted access programs.

May 2025
May 22, 2025
Anthropic activates ASL-3 deployment and security protections for Claude Opus 4
CapabilityThresholdAnthropic

Anthropic activated ASL-3 protections for Claude Opus 4 as a precaution because it could not rule out ASL-3 CBRN risk; the announcement does not cite cyber capability as the trigger. The ASL-3 security standard it describes includes more than 100 controls to protect weights, two-party authorization for weight access, and egress bandwidth controls against exfiltration.

Benchmarks and tools