Chronicle/Defense & research

Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades

DefensePaperSignificance assistant-drafted

Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.

Why it matters

Cheaper classifier guards make it more practical to run misuse safeguards on all traffic. The reported results are for CBRN safeguards; whether they carry over to cyber misuse is not shown.

Key facts

As stated in the sources, with where to find them.

  • About 1% additional compute if applied to Claude Opus 4.0 traffic, versus a 23.7% compute increase for the first-generation classifiers.Introduction; Conclusions and further research
  • 0.05% refusal rate on harmless queries over one month of deployment on Claude Sonnet 4.5 traffic, 87% lower than the original classifier system.Conclusions and further research
  • Over 1,700 cumulative red-teaming hours across 198,000 attempts found one high-risk vulnerability and no universal jailbreak.Conclusions and further research

Findings that cite this record

Key questions this bears on

Sources

Related records