Guide · AI security
Input filtering vs output validation
Input filtering checks what goes into a language model, the user's prompt and any retrieved content, and blocks or cleans it before the model runs. Output validation checks what the model produces before a user sees it or another system acts on it. Each catches failures the other misses, and OWASP and MITRE ATLAS both list the two together. Neither makes prompt injection impossible.
Last reviewed
Two checkpoints on either side of the model
An input filter runs before the model. MITRE ATLAS lists what input moderation screens user prompts for: harmful content, prompt injection attempts, jailbreak attempts, sensitive data, off-topic requests, and inputs of an unexpected length or format. ATLAS lists retrieval guardrails separately, for documents on their way into the model’s context.
Output validation runs after the model, and the name covers three different jobs.
Output moderation reads the response for harm before a user sees it: harmful content, personal data, secrets, policy violations or unsafe code. ATLAS says the check is made by classical scanners, classifiers or a dedicated reviewer model.
Output handling treats the response as untrusted data before another system uses it. OWASP calls the failure improper output handling. Its examples include model-written JavaScript that runs in a user’s browser, model-written SQL run without parameterization, and model output passed to a shell. Its advice starts with “Treat the model as any other user”, then adds encoding output for where it lands and using parameterized queries. ATLAS adds structured output validation: check the output against a schema before a downstream system consumes it.
Action checks apply when the output is a tool call. ATLAS gives tool calls their own checkpoint: validate the call, its arguments and its permissions before it runs.
Both sides can use rules, such as blocklists and schema checks, or models, such as classifiers and LLM judges. Meta’s Llama Guard is one model built to do both, classifying prompts and responses against the same risk taxonomy.
Side by side
| Input filtering | Output validation | |
|---|---|---|
| Runs | Before the model | After the model, before a user or system acts |
| Sees | The prompt, and retrieved content if filtered too | What the model actually produced |
| Stops | Harmful requests, recognised injection and jailbreak attempts, sensitive data sent to the model | Harmful or leaked content shown to users, injection reaching a browser or database, unsafe tool calls |
| Must recognise the attack | Classifiers and blocklists, yes. Length and format limits, no | Moderation, yes. Encoding, parameterized queries and schema checks, no |
| Typical failure | Misses a prompt written to look harmless; blocks legitimate text that discusses attacks | Misses output it cannot read without the prompt; in streaming, blocks after some text has gone out |
Where input filtering falls short
A classifier on the input has to recognise the attack, and the attacker writes the input. In Check Point’s PuzzleMask research, four screening models, each verified to block the plain payloads, rated every prompt they were given that hid those payloads in ordinary prose as safe. Three of them saw all 23 crafted prompts. Tested separately, a strong reasoning model with a Python code interpreter acted on the hidden instruction in 17 of 18 trials. Check Point did find an input-side fix: one added policy clause let gpt-4o-mini catch all 23 crafted prompts, at an expected cost in false positives.
Simon Willison made the wider argument in 2022. A classifier layer can itself be talked round, and against an adversary “a solution that works 99% of the time is no good”. Filters also block honest work: an injection filter on a translation app, he notes, stops it translating an article about prompt injection.
A filter on the user’s message also sees only the user’s message. OWASP describes indirect prompt injection as arriving “from external sources, such as websites or files”. That text reaches the model in a retrieved document or a tool result, so it needs a check at that point. ATLAS files those checks as retrieval guardrails and tool guardrails, separate from input moderation.
Where output validation falls short
An output check sees only the output. Anthropic’s Constitutional Classifiers paper trains its input classifier to block prompts that set up a code, where the model’s answer “cannot be understood without seeing the input”. An output-only check would read that answer as harmless.
Streaming leaves a gap too. A response shown as it is generated has to be stopped mid-stream, after some of it is already on screen. Anthropic’s output classifier scores every token to do this. In the paper’s test, the stream was not stopped early enough in 1.17% of jailbreaking attempts. During red teaming, an implementation error let testers receive up to 128 tokens after the classifier had flagged the output.
Output handling has its own limit. Encoding and parameterized queries stop injection hidden in the output, not a harmful request carried out correctly. In OWASP’s third scenario a user asks the model for a query that deletes every table, and the tables are deleted if nothing scrutinises the query.
Using both
The two checks cover each other’s blind spots. In Anthropic’s design the output classifier handles many common jailbreak strategies, and the input classifier exists to “detect and block attempts to circumvent the output classifier”. In automated tests on chemical-weapons questions, the pair refused over 95% of held-out jailbreak attempts, against 14% with no classifiers. The cost was 0.38 percentage points more refusals on a sample of production Claude.ai traffic and 23.7% more inference compute. That is one domain and one model family, so read it as evidence the combination can work, not as a benchmark to expect.
Where model output reaches a browser or a database, start with output handling. Encoding and parameterized queries do not depend on anyone recognising the attack. Willison calls escaping and parameterized queries “a known, guaranteed to work mitigation” for SQL injection, and says no one can show an AI filter is fully reliable. Add filters and moderation on top. They reduce what gets through; they do not close the hole.
Filtering and moderation do not limit what the model is allowed to do. OWASP lists least privilege and human approval for high-risk actions as separate prompt injection mitigations, alongside filtering. ATLAS adds that validation for agent components should happen “external to the AI agent”, not inside it.
Questions and answers
Should I filter LLM inputs or validate outputs?
Both, because each misses what the other catches. An input filter can stop a harmful or manipulative request before the model runs, but a classifier or blocklist has to recognise the attack. Output validation sees what the model actually produced, and its encoding and parameterized queries protect a browser or database from injection in that output whether or not anyone recognised the attack.
Does input filtering stop prompt injection?
Not reliably. OWASP says it is unclear whether any fool-proof prevention for prompt injection exists. In Check Point's PuzzleMask tests, screening models rated every crafted prompt they were given as safe, three of them across all 23; one added policy clause then let one of them catch all 23, at an expected cost in false positives. A filter that screens only the user's message also never sees instructions planted in a web page or file the model reads later.
Sources
- MITRE ATLAS, AML.M0020 Generative AI Guardrails (collection 2026.09)MITRE
- LLM05:2025 Improper Output HandlingOWASP GenAI Security Project, Nov 2024
- LLM01:2025 Prompt InjectionOWASP GenAI Security Project, Nov 2024
- You can’t solve AI security problems with more AISimon Willison, 17 Sep 2022
- PuzzleMask: Abusing Plain Prose as a Covert AI Attack VectorCheck Point Research, 10 Sep 2026
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red TeamingSharma et al., Anthropic, Jan 2025
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI ConversationsInan et al., Meta, Dec 2023