LlamaFirewall

Meta's layered guardrail framework that scans agent prompts, plans, and generated code.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Securing agents rather than chatbots is the design center of LlamaFirewall, the guardrail framework Meta runs in production and ships inside the PurpleLlama repository. Three scanner layers cover different failure modes: PromptGuard 2, an 86M-parameter classifier with a faster 22M variant, catches jailbreaks and injected instructions in user input and untrusted tool output; AlignmentCheck, an experimental chain-of-thought auditor, reads an agent's reasoning trace to detect goal hijacking mid-plan; and CodeShield, a static analyzer, flags insecure coding patterns in generated code before it executes. Scanners compose into per-conversation pipelines alongside regex and custom rules, engineered for low-latency inline use in high-throughput agent systems. The framework code is MIT licensed and installs from PyPI as llamafirewall, while the PromptGuard model weights carry the Llama community license; the small classifiers run on CPU, so no GPU is needed for the core path. An accompanying 2025 paper documents the design and benchmark results, and Meta positions it as the open backbone for agent-side defense.

Should you use LlamaFirewall?

Pick it when

Pick LlamaFirewall when you defend a tool-using agent in production and need inline, low-latency scanning of user input and tool output for prompt injection, plus checks on generated code, with core classifiers on CPU.

Look elsewhere when

Skip it for plain chatbots that need topic control or output format validation, where NeMo Guardrails or Guardrails AI fit better. The PromptGuard weights use the Llama community license, and AlignmentCheck is experimental.

Alternatives to LlamaFirewall

  • NeMo Guardrails

    Programmable Colang rails that keep conversations on approved topics and govern actions; no static analyzer for generated code like CodeShield.

  • Guardrails AI

    Validates and corrects outputs with Hub validators for PII, toxicity, and format under Apache-2.0; not built around agent prompt injection or code scanning.

  • PyRIT

    Tests your defenses offline with automated jailbreak and injection campaigns; it finds weaknesses rather than blocking them inline.

  • garak

    Scans a model endpoint for injection and leakage weaknesses before release; a testing tool, not a runtime guard.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Added
Aug 24, 2026

Tags