Chronicle/Defense & research

Meta releases LlamaFirewall guardrails with PromptGuard 2 and AlignmentCheck for agents

DefenseTool releaseSignificance assistant-drafted

Meta open-sources LlamaFirewall, combining PromptGuard 2 (a jailbreak and injection detector), AlignmentCheck (a chain-of-thought auditor for goal hijacking) and CodeShield (static analysis of generated code). On AgentDojo, Meta reports that the combination cut attack success from 17.63% to 1.75% while utility fell from 47.73% to 42.68%.

Why it matters

It is an open guardrail stack whose detector component was later shown to fail under adaptive attack.

Key facts

As stated in the sources, with where to find them.

  • AgentDojo baseline: 17.63% ASR, 47.73% utility. PromptGuard 2 86M alone: 7.53% ASR. AlignmentCheck (Llama 4 Maverick) alone: 2.89% ASR, 43.09% utility. Combined: 1.75% ASR, 42.68% utility.Section 4.3.2

Findings that cite this record

No tracked finding cites this record yet.

Sources

Related records