Chronicle/Defense & research

Anthropic introduces Constitutional Classifiers against universal jailbreaks

DefensePaperSignificance assistant-drafted

Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.

Why it matters

The classifier-guard approach was later extended to cyber misuse for Anthropic's Fable 5 safeguards.

Key facts

As stated in the sources, with where to find them.

  • On 10,000 synthetic jailbreak prompts against Claude 3.5 Sonnet (Oct 2024), success was 86% without classifiers and 4.4% with them; over-refusal rose 0.38% and compute overhead was 23.7%.Automated evaluations section
  • Prototype bug bounty: 183 active participants, over 3,000 hours, no universal jailbreak answering all ten forbidden queries.Initial red teaming section
  • Public demo (Feb 3-10, 2025): 339 jailbreakers, over 300,000 interactions, about 3,700 hours; one participant achieved a universal jailbreak across all eight levels; $55,000 paid to four winners.Demo results update

Findings that cite this record

Sources

Related records