Meta open-sources LlamaFirewall, combining PromptGuard 2 (a jailbreak and injection detector), AlignmentCheck (a chain-of-thought auditor for goal hijacking) and CodeShield (static analysis of generated code). On AgentDojo, Meta reports that the combination cut attack success from 17.63% to 1.75% while utility fell from 47.73% to 42.68%.
Why it matters
It is an open guardrail stack whose detector component was later shown to fail under adaptive attack.
Key facts
As stated in the sources, with where to find them.
- AgentDojo baseline: 17.63% ASR, 47.73% utility. PromptGuard 2 86M alone: 7.53% ASR. AlignmentCheck (Llama 4 Maverick) alone: 2.89% ASR, 43.09% utility. Combined: 1.75% ASR, 42.68% utility.Section 4.3.2
Findings that cite this record
No tracked finding cites this record yet.
Sources
Related records
Mar 24, 2025
Oct 10, 2025
Jan 31, 2025
Mar 24, 2025
Jun 9, 2026
Nov 7, 2025