<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Jailbreaks &amp; safeguards · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/jailbreaks-and-safeguards/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/jailbreaks-and-safeguards/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on jailbreaks &amp; safeguards, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-cyber-jailbreak-severity-framework-2026/</link>
<guid isPermaLink="false">event:anthropic-cyber-jailbreak-severity-framework-2026</guid>
<pubDate>Thu, 02 Jul 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response. A shared severity scale for safeguard bypasses is a precondition for proportionate government and industry responses like the June 2026 suspension.</description>
</item>
<item>
<title>US lifts export controls on Fable 5 and Mythos 5; Anthropic redeploys with new cyber classifier</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-lifts-controls-fable-5-redeployed-2026/</link>
<guid isPermaLink="false">event:us-lifts-controls-fable-5-redeployed-2026</guid>
<pubDate>Tue, 30 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic announced that export controls on Fable 5 and Mythos 5 had been lifted and that Fable 5 would be redeployed globally from July 1, 2026 with an improved safety classifier. Anthropic says the classifier blocks the technique described in an Amazon report in over 99% of cases and that CAISI researchers tested its prior and new safeguards. Mythos 5 access was restored for a set of US organizations after government approval on June 26. It shows the conditions, including government testing of safeguards, under which a suspended cyber-capable model was allowed back.</description>
</item>
<item>
<title>US export-control directive forces Anthropic to suspend Fable 5 and Mythos 5 over safeguard bypass</title>
<link>https://agentic-cyber-explorer.pages.dev/events/us-directive-suspends-fable-5-mythos-5-2026/</link>
<guid isPermaLink="false">event:us-directive-suspends-fable-5-mythos-5-2026</guid>
<pubDate>Fri, 12 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>Anthropic said the US government issued an export control directive, citing national security authorities, barring access to Fable 5 and Mythos 5 by foreign nationals, after officials said they had found a way to jailbreak Fable 5's safeguards. Anthropic said the net effect was that it had to disable both models for all customers to comply, while other Claude models stayed available. Anthropic disputed the rationale, arguing the demonstrated vulnerabilities were minor and that the standard applied industry-wide would halt new frontier deployments. It is a case of a government using export controls to pull a deployed frontier model over a cyber-safeguard bypass.</description>
</item>
<item>
<title>NIST scientist argues no finite guardrail set is robust to adversarial prompts, urges continuous updates</title>
<link>https://agentic-cyber-explorer.pages.dev/events/nist-no-finite-guardrails-continuous-monitoring-2026/</link>
<guid isPermaLink="false">event:nist-no-finite-guardrails-continuous-monitoring-2026</guid>
<pubDate>Tue, 09 Jun 2026 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>NIST announced a paper by Apostol Vassilev in IEEE Security &amp; Privacy arguing, by extension of Gödel's incompleteness results, that no finite set of guardrails can be universally robust against adversarial prompts. NIST recommends a continuous monitor-and-update model: ongoing red teaming, continuous guardrail updates, and operational resilience to limit impact and recover. It gives US government backing to treating jailbreak and injection defense for agents as an ongoing operational process rather than a certifiable property.</description>
</item>
<item>
<title>OpenAI releases IH-Challenge RL dataset and reports instruction hierarchy gains on injection benchmarks</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-ih-challenge-dataset-2026/</link>
<guid isPermaLink="false">event:openai-ih-challenge-dataset-2026</guid>
<pubDate>Tue, 10 Mar 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>OpenAI describes IH-Challenge, a reinforcement learning dataset of simple, programmatically graded conflicts between higher- and lower-privilege instructions designed to avoid shortcuts such as over-refusal. A GPT-5 Mini variant trained on it (GPT-5 Mini-R) improved on instruction-hierarchy benchmarks and on CyberSecEval 2 and an internal prompt injection benchmark, with little capability loss; the dataset is publicly released. It is an open training resource for model-level prompt injection robustness from a frontier lab.</description>
</item>
<item>
<title>Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-next-gen-constitutional-classifiers-2026/</link>
<guid isPermaLink="false">event:anthropic-next-gen-constitutional-classifiers-2026</guid>
<pubDate>Fri, 09 Jan 2026 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours. Cheaper classifier guards make it more practical to run misuse safeguards on all traffic. The reported results are for CBRN safeguards; whether they carry over to cyber misuse is not shown.</description>
</item>
<item>
<title>Anthropic disrupts a state-sponsored espionage campaign it says was largely executed by Claude Code</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-ai-orchestrated-espionage-gtg-1002-2025/</link>
<guid isPermaLink="false">event:anthropic-ai-orchestrated-espionage-gtg-1002-2025</guid>
<pubDate>Thu, 13 Nov 2025 12:00:00 GMT</pubDate>
<category>Attacks &amp; incidents</category>
<description>Anthropic reports that in mid-September 2025 a group it assesses with high confidence to be Chinese state-sponsored used Claude Code inside an attack framework to attempt intrusions into about thirty organizations, succeeding in a small number. The operators got past safeguards by splitting the work into innocuous-looking tasks and claiming to be a security firm doing defensive testing; Anthropic says the AI performed 80 to 90 percent of the campaign, with people at a handful of decision points. It is Anthropic's account of an AI agent executing most of a state espionage operation against real targets, which it tracks as GTG-1002.</description>
</item>
<item>
<title>Google reports malware that queries LLMs during execution, including APT28's PROMPTSTEAL</title>
<link>https://agentic-cyber-explorer.pages.dev/events/gtig-ai-threat-tracker-llm-querying-malware-2025/</link>
<guid isPermaLink="false">event:gtig-ai-threat-tracker-llm-querying-malware-2025</guid>
<pubDate>Wed, 05 Nov 2025 12:00:00 GMT</pubDate>
<category>Attacks &amp; incidents</category>
<description>Google Threat Intelligence Group's AI Threat Tracker says adversaries moved beyond productivity uses in 2025 and began deploying malware that calls LLMs mid-execution, such as PROMPTFLUX, which asks Gemini to rewrite its own code, and PROMPTSTEAL, which queries a hosted open model for commands. GTIG attributes PROMPTSTEAL to Russia's APT28 in operations against Ukraine, and also reports actors posing as CTF players or researchers to get past safeguards and a maturing underground market for AI tools. It is Google's evidence that malware using models at runtime had reached a state operation, after CERT-UA's earlier report of the same malware, and it replaced Google's own productivity-only picture.</description>
</item>
<item>
<title>'The Attacker Moves Second': adaptive attacks bypass 12 published jailbreak and injection defenses</title>
<link>https://agentic-cyber-explorer.pages.dev/events/attacker-moves-second-adaptive-attacks-2025/</link>
<guid isPermaLink="false">event:attacker-moves-second-adaptive-attacks-2025</guid>
<pubDate>Fri, 10 Oct 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Nasr, Carlini, Tramèr and 11 co-authors apply gradient, reinforcement learning, search and human red-teaming attacks to 12 published defenses. Most defenses originally reported near-zero attack success, but the adaptive attacks exceed 90% success against most, and human red-teamers succeeded on every challenge in the subset of defenses they were given. It is the central evidence that static-benchmark robustness claims for prompt injection defenses do not hold against adaptive attackers.</description>
</item>
<item>
<title>Anthropic introduces Constitutional Classifiers against universal jailbreaks</title>
<link>https://agentic-cyber-explorer.pages.dev/events/anthropic-constitutional-classifiers-2025/</link>
<guid isPermaLink="false">event:anthropic-constitutional-classifiers-2025</guid>
<pubDate>Mon, 03 Feb 2025 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak. The classifier-guard approach was later extended to cyber misuse for Anthropic's Fable 5 safeguards.</description>
</item>
<item>
<title>Google finds government-backed hackers using Gemini for support tasks, not novel capabilities</title>
<link>https://agentic-cyber-explorer.pages.dev/events/gtig-adversarial-misuse-gemini-2025/</link>
<guid isPermaLink="false">event:gtig-adversarial-misuse-gemini-2025</guid>
<pubDate>Wed, 29 Jan 2025 12:00:00 GMT</pubDate>
<category>Attacks &amp; incidents</category>
<description>Google Threat Intelligence Group analyzed how government-backed hacking and information-operations actors used the Gemini web app. It reports use for research, troubleshooting code and producing content across several attack phases, with Iranian actors the heaviest users, and says it saw productivity gains but no novel capabilities; requests for clearly malicious help drew safety responses. An independent provider reached the same conclusion as Microsoft and OpenAI a year earlier, shortly before reports of agentic misuse began later in 2025.</description>
</item>
<item>
<title>NIST second draft of AI 800-1 on dual-use foundation model misuse adds cybersecurity appendix</title>
<link>https://agentic-cyber-explorer.pages.dev/events/nist-ai-800-1-second-draft-cyber-misuse-2025/</link>
<guid isPermaLink="false">event:nist-ai-800-1-second-draft-cyber-misuse-2025</guid>
<pubDate>Wed, 15 Jan 2025 12:00:00 GMT</pubDate>
<category>Policy &amp; standards</category>
<description>NIST's AI Safety Institute released a second public draft of NIST AI 800-1, voluntary guidelines for managing misuse risk from dual-use foundation models across the lifecycle. NIST says the draft adds detailed evaluation approaches, a marginal-risk framework, and an extensive appendix on cybersecurity misuse risk, and covers both closed and open model developers. It was the main US government draft practice for measuring and mitigating cyber misuse of frontier models before the 2025 policy shift.</description>
</item>
<item>
<title>OpenAI trains models to prioritize privileged instructions via an instruction hierarchy</title>
<link>https://agentic-cyber-explorer.pages.dev/events/openai-instruction-hierarchy-2024/</link>
<guid isPermaLink="false">event:openai-instruction-hierarchy-2024</guid>
<pubDate>Fri, 19 Apr 2024 12:00:00 GMT</pubDate>
<category>Defense &amp; research</category>
<description>Wallace and co-authors at OpenAI argue that models treat system prompts and untrusted inputs with equal priority and propose an explicit instruction hierarchy that tells the model which instructions to follow when they conflict. Applied to GPT-3.5, they report large robustness gains against attack types not seen in training with minimal capability loss. The instruction hierarchy became OpenAI's stated foundation for prompt-injection robustness in later agent products.</description>
</item>
</channel>
</rss>
