Constitutional Classifiers

Input and output classifiers trained from a written constitution to block jailbreaks.

Records citing Constitutional Classifiers

Jan 9, 2026
Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades
DefensePaperAnthropic

Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.

Feb 3, 2025
Anthropic introduces Constitutional Classifiers against universal jailbreaks
DefensePaperAnthropic

Anthropic describes input and output classifiers trained on synthetic data generated from a natural-language constitution of allowed and disallowed content, targeted at chemical-weapons style queries. In automated testing on Claude 3.5 Sonnet, jailbreak success fell from 86% to 4.4%, and a prior bug bounty found no universal jailbreak; a public demo in February 2025 did yield one universal jailbreak.