Anthropic describes Constitutional Classifiers++, a cascade in which a cheap linear probe on model activations screens all traffic and escalates flagged exchanges to a probe-classifier ensemble. It reports roughly 1% added compute if applied to Claude Opus 4.0 traffic (the first generation added 23.7%) and a 0.05% refusal rate on harmless queries over one month of Claude Sonnet 4.5 traffic. Red-teamers found no universal jailbreak in over 1,700 hours.
Week of Jan 5–11, 2026
Defense & research
Anthropic reports that Pacific Northwest National Laboratory built a scaffold around Claude Sonnet 4 to automate adversary emulation against a high-fidelity cyber-physical model of a water treatment plant used for CISA. PNNL estimates attack reconstruction took three hours instead of multiple weeks; in one run the model switched to a different known privilege-escalation technique when a provided tool failed.
Policy & standards
CAISI published a Federal Register request for information on practices for measuring and improving the security of AI agent systems, citing hijacking, backdoors and indirect prompt injection. It asks about model-level, system-level and human-oversight controls, assessment methods, and ways to limit, modify and monitor deployment environments.