Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.
Why it matters
The classifier-guard approach was later extended to cyber misuse for Anthropic's Fable 5 safeguards.
Key facts
As stated in the sources, with where to find them.
- On 10,000 synthetic jailbreak prompts against Claude 3.5 Sonnet (Oct 2024), success was 86% without classifiers and 4.4% with them; over-refusal rose 0.38% and compute overhead was 23.7%.Automated evaluations section
- Prototype bug bounty: 183 active participants, over 3,000 hours, no universal jailbreak answering all ten forbidden queries.Initial red teaming section
- Public demo (Feb 3-10, 2025): 339 jailbreakers, over 300,000 interactions, about 3,700 hours; one participant achieved a universal jailbreak across all eight levels; $55,000 paid to four winners.Demo results update
Findings that cite this record
Sources
Related records
Jan 9, 2026
Jul 2, 2026
Jun 30, 2026
Jun 12, 2026
Jan 15, 2025
Jun 17, 2025