Findings/classifiers-cut-jailbreaks-not-to-zero

Anthropic reports that classifier guards cut automated jailbreak success from 86% to 4.4%; a 2025 public demo still yielded one universal jailbreak, though none has been reported against its 2026 successor.

Reportedreported2 evidence records from 1 independent sourceassistant-drafted
Scope: what this does not show

Vendor-reported results for one lab's safeguards, built for chemical-weapons and other CBRN content rather than cyber misuse.

Reported: Stated by one source and not yet corroborated or challenged.

Evidence

Status history

  1. 2025-02-03ReportedAnthropic Constitutional Classifiers. · record