Scope: what this does not show
Vendor-reported results for one lab's safeguards, built for chemical-weapons and other CBRN content rather than cyber misuse.
Reported: Stated by one source and not yet corroborated or challenged.
Evidence
Feb 3, 2025
Anthropic introduces Constitutional Classifiers against universal jailbreaks
Automated jailbreak success on Claude 3.5 Sonnet fell from 86% to 4.4%; in the Feb 3-10, 2025 public demo one participant found a universal jailbreak.
Jan 9, 2026
Anthropic's next-generation Constitutional Classifiers cut overhead to about 1% using probe cascades
Over 1,700 red-teaming hours and 198,000 attempts found no universal jailbreak; about 1% compute overhead.
Status history
- 2025-02-03ReportedAnthropic Constitutional Classifiers. · record