Topics/Governance

Jailbreaks & safeguards

Model-level safeguards against cyber misuse and attempts to bypass them.

13 records2 findings0 openings4 benchmarks and toolsLatest record
RangeLanes
12 of 13 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
20252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

12 records · newest first
Jul 2026
Jul 2, 2026
Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners
PolicyFrameworkAnthropic

Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response.

Jun 2026
Jun 30, 2026
US lifts export controls on Fable 5 and Mythos 5; Anthropic redeploys with new cyber classifier
PolicyRegulationAnthropic, US Department of Commerce, US Center for AI Standards and Innovation

Anthropic announced that export controls on Fable 5 and Mythos 5 had been lifted and that Fable 5 would be redeployed globally from July 1, 2026 with an improved safety classifier. Anthropic says the classifier blocks the technique described in an Amazon report in over 99% of cases and that CAISI researchers tested its prior and new safeguards. Mythos 5 access was restored for a set of US organizations after government approval on June 26.

Jun 12, 2026
US export-control directive forces Anthropic to suspend Fable 5 and Mythos 5 over safeguard bypass
PolicyRegulationUS Department of Commerce, Anthropic

Anthropic said the US government issued an export control directive, citing national security authorities, barring access to Fable 5 and Mythos 5 by foreign nationals, after officials said they had found a way to jailbreak Fable 5's safeguards. Anthropic said the net effect was that it had to disable both models for all customers to comply, while other Claude models stayed available. Anthropic disputed the rationale, arguing the demonstrated vulnerabilities were minor and that the standard applied industry-wide would halt new frontier deployments.

Jun 9, 2026
NIST scientist argues no finite guardrail set is robust to adversarial prompts, urges continuous updates
PolicyGuidanceNIST

NIST announced a paper by Apostol Vassilev in IEEE Security & Privacy arguing, by extension of Gödel's incompleteness results, that no finite set of guardrails can be universally robust against adversarial prompts. NIST recommends a continuous monitor-and-update model: ongoing red teaming, continuous guardrail updates, and operational resilience to limit impact and recover.

Mar 2026
Mar 10, 2026
OpenAI releases IH-Challenge RL dataset and reports instruction hierarchy gains on injection benchmarks
DefenseDatasetOpenAI

OpenAI describes IH-Challenge, a reinforcement learning dataset of simple, programmatically graded conflicts between higher- and lower-privilege instructions designed to avoid shortcuts such as over-refusal. A GPT-5 Mini variant trained on it (GPT-5 Mini-R) improved on instruction-hierarchy benchmarks and on CyberSecEval 2 and an internal prompt injection benchmark, with little capability loss; the dataset is publicly released.

Jan 2026
Jan 9, 2026
Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades
DefensePaperAnthropic

Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.

Nov 2025
Nov 13, 2025
Anthropic disrupts a state-sponsored espionage campaign it says was largely executed by Claude Code
AttackMisuse reportAnthropic, GTG-1002

Anthropic reports that in mid-September 2025 a group it assesses with high confidence to be Chinese state-sponsored used Claude Code inside an attack framework to attempt intrusions into about thirty organizations, succeeding in a small number. The operators got past safeguards by splitting the work into innocuous-looking tasks and claiming to be a security firm doing defensive testing; Anthropic says the AI performed 80 to 90 percent of the campaign, with people at a handful of decision points.

Nov 5, 2025
Google reports malware that queries LLMs during execution, including APT28's PROMPTSTEAL
AttackMisuse reportGoogle Threat Intelligence Group, APT28

Google Threat Intelligence Group's AI Threat Tracker says adversaries moved beyond productivity uses in 2025 and began deploying malware that calls LLMs mid-execution, such as PROMPTFLUX, which asks Gemini to rewrite its own code, and PROMPTSTEAL, which queries a hosted open model for commands. GTIG attributes PROMPTSTEAL to Russia's APT28 in operations against Ukraine, and also reports actors posing as CTF players or researchers to get past safeguards and a maturing underground market for AI tools.

Oct 2025
Feb 2025
Feb 3, 2025
Anthropic introduces Constitutional Classifiers against universal jailbreaks
DefensePaperAnthropic

Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.

Jan 2025
Jan 29, 2025
Google finds government-backed hackers using Gemini for support tasks, not novel capabilities
AttackMisuse reportGoogle Threat Intelligence Group, Google

Google Threat Intelligence Group analyzed how government-backed hacking and information-operations actors used the Gemini web app. It reports use for research, troubleshooting code and producing content across several attack phases, with Iranian actors the heaviest users, and says it saw productivity gains but no novel capabilities; requests for clearly malicious help drew safety responses.

Jan 15, 2025
NIST second draft of AI 800-1 on dual-use foundation model misuse adds cybersecurity appendix
PolicyGuidanceNIST, US Center for AI Standards and Innovation

NIST's AI Safety Institute released a second public draft of NIST AI 800-1, voluntary guidelines for managing misuse risk from dual-use foundation models across the lifecycle. NIST says the draft adds detailed evaluation approaches, a marginal-risk framework, and an extensive appendix on cybersecurity misuse risk, and covers both closed and open model developers.

Findings

Benchmarks and tools