Methods/Attack technique

Jailbreaking safeguards

Inputs crafted to make a model ignore its safety training or safeguard classifiers, for example to obtain help it would normally refuse.

10 records3 attack1 capability4 defense2 policy2 findings (1 measured)First recorded 2025-01assistant-drafted

How it works

Attackers search for phrasings, role-play framings, or multi-turn sequences that the model's training and classifiers fail to recognize as disallowed.

What we know

2 reported

Records over time

RangeLanes
10 of 10 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
20252026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

10 records · newest first
Jul 2026
Jul 2, 2026
Anthropic proposes Cyber Jailbreak Severity scale with Glasswing partners
PolicyFrameworkAnthropic

Anthropic published an early-draft Cyber Jailbreak Severity framework, developed with Project Glasswing partners, to score cyber jailbreaks on capability gain, breadth, ease of weaponization and discoverability, mapped to five levels from CJS-0 to CJS-4. It also described Fable 5's cyber classifier tiers, which block prohibited and high-risk dual-use requests such as exploit development while allowing defensive work like patching and incident response.

Jun 2026
Jun 12, 2026
US export-control directive forces Anthropic to suspend Fable 5 and Mythos 5 over safeguard bypass
PolicyRegulationUS Department of Commerce, Anthropic

Anthropic said the US government issued an export control directive, citing national security authorities, barring access to Fable 5 and Mythos 5 by foreign nationals, after officials said they had found a way to jailbreak Fable 5's safeguards. Anthropic said the net effect was that it had to disable both models for all customers to comply, while other Claude models stayed available. Anthropic disputed the rationale, arguing the demonstrated vulnerabilities were minor and that the standard applied industry-wide would halt new frontier deployments.

Jan 2026
Jan 9, 2026
Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades
DefensePaperAnthropic

Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.

Nov 2025
Nov 13, 2025
Anthropic disrupts a state-sponsored espionage campaign it says was largely executed by Claude Code
AttackMisuse reportAnthropic, GTG-1002

Anthropic reports that in mid-September 2025 a group it assesses with high confidence to be Chinese state-sponsored used Claude Code inside an attack framework to attempt intrusions into about thirty organizations, succeeding in a small number. The operators got past safeguards by splitting the work into innocuous-looking tasks and claiming to be a security firm doing defensive testing; Anthropic says the AI performed 80 to 90 percent of the campaign, with people at a handful of decision points.

Nov 5, 2025
Google reports malware that queries LLMs during execution, including APT28's PROMPTSTEAL
AttackMisuse reportGoogle Threat Intelligence Group, APT28

Google Threat Intelligence Group's AI Threat Tracker says adversaries moved beyond productivity uses in 2025 and began deploying malware that calls LLMs mid-execution, such as PROMPTFLUX, which asks Gemini to rewrite its own code, and PROMPTSTEAL, which queries a hosted open model for commands. GTIG attributes PROMPTSTEAL to Russia's APT28 in operations against Ukraine, and also reports actors posing as CTF players or researchers to get past safeguards and a maturing underground market for AI tools.

Oct 2025
Sep 2025
Sep 30, 2025
CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack
CapabilityEvaluation reportUS Center for AI Standards and Innovation, NIST, DeepSeek

NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.

May 2025
May 6, 2025
Meta releases LlamaFirewall guardrails with PromptGuard 2 and AlignmentCheck for agents
DefenseTool releaseMeta

Meta open-sources LlamaFirewall, combining PromptGuard 2 (a jailbreak and injection detector), AlignmentCheck (a chain-of-thought auditor for goal hijacking) and CodeShield (static analysis of generated code). On AgentDojo, Meta reports that the combination cut attack success from 17.63% to 1.75% while utility fell from 47.73% to 42.68%.

Feb 2025
Feb 3, 2025
Anthropic introduces Constitutional Classifiers against universal jailbreaks
DefensePaperAnthropic

Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.

Jan 2025
Jan 29, 2025
Google finds government-backed hackers using Gemini for support tasks, not novel capabilities
AttackMisuse reportGoogle Threat Intelligence Group, Google

Google Threat Intelligence Group analyzed how government-backed hacking and information-operations actors used the Gemini web app. It reports use for research, troubleshooting code and producing content across several attack phases, with Iranian actors the heaviest users, and says it saw productivity gains but no novel capabilities; requests for clearly malicious help drew safety responses.

All records