Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.
Why it matters
Cheaper classifier guards make it more practical to run misuse safeguards on all traffic. The reported results are for CBRN safeguards; whether they carry over to cyber misuse is not shown.
Key facts
As stated in the sources, with where to find them.
- About 1% additional compute if applied to Claude Opus 4.0 traffic, versus a 23.7% compute increase for the first-generation classifiers.Introduction; Conclusions and further research
- 0.05% refusal rate on harmless queries over one month of deployment on Claude Sonnet 4.5 traffic, 87% lower than the original classifier system.Conclusions and further research
- Over 1,700 cumulative red-teaming hours across 198,000 attempts found one high-risk vulnerability and no universal jailbreak.Conclusions and further research
Findings that cite this record
Key questions this bears on
- Is AI shifting the balance between finding and fixing vulnerabilities?Discovery is ahead. AI finds real vulnerabilities faster than they are fixed, and simple checks overstate how often AI patches work.
Sources
Related records
Nov 24, 2025
Feb 3, 2025
Feb 5, 2026
Jul 2, 2026
Jun 30, 2026
Jun 17, 2025