What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

AI guardrails

AI guardrails are safety controls placed between users, tools, and a generative AI model, checking prompts, retrieved content, model output, and agent actions before they are accepted, executed, or shown to a user. MITRE ATLAS lists them as a mitigation against prompt injection, jailbreaks, data leakage, and several other generative AI attacks.

Last reviewed

Key points

  • MITRE ATLAS defines guardrails as controls placed between users, tools, and the model that evaluate prompts, retrieved context, outputs, and agent actions before they take effect, rather than changing the model itself.
  • ATLAS's 2026.08 collection maps guardrails to 23 of its listed attack techniques, more than any mitigation except AI red teaming, spanning prompt injection, jailbreaks, system prompt extraction, LLM data leakage, and RAG poisoning.
  • Guardrails can be rule-based, such as filters, allowlists, and schema checks, or AI-based, such as classifier and LLM reviewer models, applied at input, retrieval, tool-call, and output checkpoints.
  • Guardrails sit around a model and need no retraining. Model alignment changes the model itself through training, and ATLAS says the two should be used together, not as substitutes.
  • Guardrails can be beaten. Datadog warns sophisticated attackers can model and circumvent static filters, and ATLAS itself says guardrails need continuous evaluation and red-teaming as attacks evolve.

A generative AI model cannot fully police its own inputs and outputs. Guardrails are the layer built around it to try: controls that sit between a user or tool and the model, checking what goes in and what comes out before either takes effect.

How guardrails work

MITRE ATLAS places guardrail checks at several points in a generative AI workflow: screening a prompt before it reaches the model, filtering documents a retrieval system pulls into context, validating a tool call and its arguments before it runs, and reviewing the model’s output before a user sees it.

The checks split two ways. Rule-based guardrails are filters, allowlists, blocklists, and schema validation — deterministic and easy to audit. AI-based guardrails are classifiers, LLM reviewer models, and groundedness checks that compare an answer against a trusted source, catching what a fixed rule cannot name in advance. NVIDIA’s NeMo Guardrails library splits its checks the same way ATLAS does: input, dialog, retrieval, execution, and output rails, one per stage of the workflow.

Why guardrails matter

Guardrails cover more ground than most single mitigations. In ATLAS’s 2026.08 collection, AML.M0020 maps to 23 of the framework’s listed techniques, more than any mitigation except AI red teaming, spanning prompt injection, jailbreaking, system prompt extraction, LLM data leakage, and RAG poisoning.

They are also a runtime layer, not a training one. Model alignment changes what a model does by retraining it; guardrails change what a deployment allows without touching the model, which is why ATLAS says to use both together. That runtime position is also the limit: Datadog warns that sophisticated attackers can model and circumvent static filters, so ATLAS calls for guardrails to be continuously evaluated and red-teamed as the attacks against them evolve, not installed once and left alone.

In practice

PuzzleMask shows how an LLM that screens prompts can be slipped past. It hides a forbidden instruction in ordinary prose. In Check Point’s tests, three screening models rated all 23 crafted prompts safe, and the stronger model behind the screen then acted on the hidden instruction in 17 of 18 trials. Checking outputs and actions as well as inputs is one of the defences Check Point suggests.

Questions and answers

Are AI guardrails the same as model alignment?

No. MITRE ATLAS files them as separate mitigations that act at different times. Model alignment retrains the model itself so it behaves differently by default. Guardrails sit outside the model and check what goes in and out at deployment time, needing no retraining to change. ATLAS says the two should be used in combination rather than as substitutes for each other.

Can AI guardrails be bypassed?

Yes. Datadog warns that sophisticated attackers can model and circumvent static filters, and ATLAS itself says guardrails should be continuously evaluated, red-teamed, and updated as adversarial techniques evolve. A guardrail is maintained, not installed once.

Sources

  1. MITRE ATLAS, AML.M0020 Generative AI Guardrails (collection 2026.08)MITRE
  2. NeMo Guardrails Library, Rail TypesNVIDIA
  3. LLM guardrails: Best practices for deploying LLM apps securelyDatadog

Guides that use this term

In the news