<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Access controls · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/access-controls/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/access-controls/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on access controls, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>Google releases Gemini 3.8 Flash Cyber for trusted defenders, emphasizing automated patching</title>
<link>https://agentic-cyber-explorer.pages.dev/events/google-gemini-3-8-flash-cyber-2026/</link>
<guid isPermaLink="false">event:google-gemini-3-8-flash-cyber-2026</guid>
<pubDate>Wed, 02 Sep 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Google introduced Gemini 3.8 Flash Cyber, a cybersecurity-tuned model with more permissive cyber mitigations, available only to trusted defenders through a new Fairwind Program. Google says it prioritized vulnerability fixing over exploitation and reports 47.2% pass@1 on Collinear's CWE-Bench patching benchmark, over 70% on an internal 20-language discovery benchmark, and 2.6 times more correct Chrome patches than larger commercial models. It is a gated, defense-oriented model release that foregrounds patching metrics rather than offensive capability.</description>
</item>
<item>
<title>European Commission presents EU Action Plan on Cybersecurity and Artificial Intelligence</title>
<link>https://agentic-cyber-explorer.pages.dev/events/eu-action-plan-cybersecurity-ai-2026/</link>
<guid isPermaLink="false">event:eu-action-plan-cybersecurity-ai-2026</guid>
<pubDate>Tue, 07 Jul 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The Commission presented an action plan responding to advanced AI models that can both improve and undermine cybersecurity. It plans an EU capacity to evaluate AI models, a European blueprint for structured access to advanced AI capabilities developed with ENISA, a secure ENISA-JRC platform to test AI for cybersecurity, AI-assisted vulnerability fixing, and a campaign to secure critical open-source software. ENISA published its own recommendations for the frontier AI era the same day. It is a dedicated EU policy response to frontier AI cyber capability, including structured access for defenders.</description>
</item>
<item>
<title>Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-cyber-jailbreak-severity-framework-2026/</link>
<guid isPermaLink="false">event:anthropic-cyber-jailbreak-severity-framework-2026</guid>
<pubDate>Thu, 02 Jul 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response. A shared severity scale for safeguard bypasses is a precondition for proportionate government and industry responses like the June 2026 suspension.</description>
</item>
<item>
<title>US lifts export controls on Fable 5 and Mythos 5; Anthropic redeploys with new cyber classifier</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-lifts-controls-fable-5-redeployed-2026/</link>
<guid isPermaLink="false">event:us-lifts-controls-fable-5-redeployed-2026</guid>
<pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic announced that export controls on Fable 5 and Mythos 5 had been lifted and that Fable 5 would be redeployed globally from July 1, 2026 with an improved safety classifier. Anthropic says the classifier blocks the technique described in an Amazon report in over 99% of cases and that CAISI researchers tested its prior and new safeguards. Mythos 5 access was restored for a set of US organizations after government approval on June 26. It shows the conditions, including government testing of safeguards, under which a suspended cyber-capable model was allowed back.</description>
</item>
<item>
<title>US export-control directive forces Anthropic to suspend Fable 5 and Mythos 5 over safeguard bypass</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-directive-suspends-fable-5-mythos-5-2026/</link>
<guid isPermaLink="false">event:us-directive-suspends-fable-5-mythos-5-2026</guid>
<pubDate>Fri, 12 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic said the US government issued an export control directive, citing national security authorities, barring access to Fable 5 and Mythos 5 by foreign nationals, after officials said they had found a way to jailbreak Fable 5's safeguards. Anthropic said the net effect was that it had to disable both models for all customers to comply, while other Claude models stayed available. Anthropic disputed the rationale, arguing the demonstrated vulnerabilities were minor and that the standard applied industry-wide would halt new frontier deployments. It is a case of a government using export controls to pull a deployed frontier model over a cyber-safeguard bypass.</description>
</item>
<item>
<title>Executive Order 14409 creates classified cyber benchmarking for covered frontier models and a clearinghouse</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-eo-14409-frontier-ai-cyber-benchmarking-2026/</link>
<guid isPermaLink="false">event:us-eo-14409-frontier-ai-cyber-benchmarking-2026</guid>
<pubDate>Tue, 02 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Executive Order 14409 directs Treasury, NSA and CISA to develop a classified benchmarking process to assess advanced cyber capabilities of AI models and designate covered frontier models, with a voluntary framework for pre-release government and trusted-partner access. It also orders an AI cybersecurity clearinghouse to coordinate vulnerability scanning, validation and remediation with industry, and states it does not create mandatory licensing or pre-clearance. It is a US mechanism that designates models by cyber capability and gives the government early access before release to other trusted partners.</description>
</item>
<item>
<title>Glasswing update: over 10,000 high-severity bugs found, but only 75 of 530 disclosed OSS bugs patched</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-glasswing-initial-update-2026/</link>
<guid isPermaLink="false">event:anthropic-glasswing-initial-update-2026</guid>
<pubDate>Fri, 22 May 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic reports that about 50 Glasswing partners used Claude Mythos Preview to find more than ten thousand high- or critical-severity vulnerabilities, and that its own scan of over 1,000 open-source projects produced 6,202 model-estimated high/critical findings. Of 1,752 assessed, mostly by six independent firms, 90.6% were true positives; Anthropic estimates 530 high/critical bugs disclosed, of which 75 were patched, and says triage and patching capacity, not discovery, is the bottleneck. It gives rare pipeline-level numbers showing AI vulnerability discovery outpacing the human capacity to verify, disclose and fix.</description>
</item>
<item>
<title>Anthropic launches Project Glasswing to give defenders early access to Claude Mythos Preview</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-project-glasswing-2026/</link>
<guid isPermaLink="false">event:anthropic-project-glasswing-2026</guid>
<pubDate>Tue, 07 Apr 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic launched Project Glasswing with AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks to use the unreleased Claude Mythos Preview for defensive security work, extending access to over 40 more organizations that maintain critical software. Anthropic committed up to $100M in usage credits and $4M in donations to open-source security groups, and reports Mythos Preview found thousands of high-severity vulnerabilities, including in every major operating system and browser. It is a large, restricted-access defensive deployment of a model its developer does not plan to make generally available, pending safeguards for Mythos-class models.</description>
</item>
<item>
<title>Frontier Model Forum report sets out shared cyber thresholds for frontier AI safety frameworks</title>
<link>https://agentic-cyber-explorer.pages.dev/events/fmf-managing-advanced-cyber-risks-frameworks-2026/</link>
<guid isPermaLink="false">event:fmf-managing-advanced-cyber-risks-frameworks-2026</guid>
<pubDate>Fri, 13 Feb 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>The Frontier Model Forum published a technical report on managing advanced cyber risks within frontier AI safety frameworks. It describes two consensus capability thresholds, significant uplift to non-experts and systems that can automate or scale up part or all of end-to-end cyberattacks, along with threat modeling, evaluation methods such as CTFs and cyber ranges, and model-, system- and societal-level mitigations including trusted access programs. It is the closest thing to an industry consensus definition of when a model's cyber capability should trigger stronger controls.</description>
</item>
<item>
<title>OpenAI adds Lockdown Mode and Elevated Risk labels to ChatGPT to limit prompt injection exfiltration</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-lockdown-mode-elevated-risk-2026/</link>
<guid isPermaLink="false">event:openai-lockdown-mode-elevated-risk-2026</guid>
<pubDate>Fri, 13 Feb 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>OpenAI introduced Lockdown Mode, an optional setting that deterministically disables or limits capabilities an attacker could exploit through prompt injection, such as live web access, image support in responses, Deep Research, Agent Mode, live connectors and file downloads. Elevated Risk labels flag network-related features in ChatGPT, Atlas and Codex that carry extra risk. Lockdown Mode first launched for enterprise-type plans, and a June 4, 2026 update says it is rolling out to personal and self-serve Business accounts. A major vendor chose to offer capability removal, not only detection, as the stronger control for high-risk users.</description>
</item>
<item>
<title>Anthropic activates ASL-3 deployment and security protections for Claude Opus 4</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-asl3-activation-claude-opus-4-2025/</link>
<guid isPermaLink="false">event:anthropic-asl3-activation-claude-opus-4-2025</guid>
<pubDate>Thu, 22 May 2025 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>Anthropic activated ASL-3 protections for Claude Opus 4 as a precaution because it could not rule out ASL-3 CBRN risk; the announcement does not cite cyber capability as the trigger. The ASL-3 security standard it describes includes more than 100 controls to protect weights, two-party authorization for weight access, and egress bandwidth controls against exfiltration. It was a public activation of a higher safety level under a lab framework, and its egress and access controls are defenses against cyber theft of model weights.</description>
</item>
<item>
<title>Anthropic RSP v2 lists cyber operations as a capability under ongoing assessment, not a threshold</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-rsp-v2-cyber-operations-assessment-2024/</link>
<guid isPermaLink="false">event:anthropic-rsp-v2-cyber-operations-assessment-2024</guid>
<pubDate>Tue, 15 Oct 2024 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic's Responsible Scaling Policy version 2.0, effective October 15, 2024, and its 2.x revisions list cyber operations among capabilities requiring ongoing assessment rather than as a formal capability threshold with required safeguards. The policy says Anthropic would consult cyber experts, consider tiered access controls or phased deployment for models with advanced cyber capabilities, and run pre- or post-deployment testing. It shows how one major lab treated autonomous exploit development as a monitored risk without committing to a cyber-specific ASL trigger during 2024-2025.</description>
</item>
</channel>
</rss>
