What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Prompt injection mitigations

Prompt injection mitigations are the controls that limit the harm prompt injection can do to a system built on a language model. Because no reliable way to stop a model being fooled exists, the controls that hold assume the model will be fooled and constrain what it can do, reach and remember.

Last reviewed

Key points

  • OWASP states that no reliable way to prevent prompt injection exists today, so defence must limit what a fooled model can do rather than try to catch every attack.
  • Controls that limit damage after an injection succeeds are what OWASP expects to survive a probing attacker; controls that lower its odds weaken as attackers adapt.
  • For agents, OWASP calls least privilege and capability budgeting "load-bearing". Keep credentials in application code, cap what one agent can reach, and put a human before privileged, irreversible or externally visible actions.
  • Input filters lower the odds but do not stop an attack. The NCSC also advises monitoring model inputs, outputs and tool calls, where failed calls could reveal an attack early.

Which controls hold

OWASP says no single prompt injection control is sufficient, and separates two effects. Some controls “reduce injection success and are expected to degrade against adaptive attackers”. Others limit the damage once an injection succeeds, and “are what survive against attackers who can probe the system”. This page sorts most of OWASP’s controls by that line.

Controls that limit the damage:

  • Least privilege per operation. Credentials and state-changing actions stay in application code, and fixed rules, not a model, re-check each privileged call. The NCSC quotes a rule of thumb from Simon Willison and Baibhav Bista: a model reading a party’s information gets that party’s privileges.
  • A budget on what one agent can hold, with the Agents Rule of Two as OWASP’s floor.
  • A human in the loop before privileged, irreversible or externally visible actions, shown the exact action. Approval fatigue wears this down.
  • Memory writes treated as privileged, so an injection is harder to carry into later sessions.

Controls that lower the odds: input filters, stripping invisible Unicode, marking retrieved content as data, and stating the model’s role in the system prompt, which OWASP calls “a partial control only”.

Why limiting damage comes first

Odds-lowering defences weaken against attackers who adapt. Nasr et al. bypassed 12 recent defences against jailbreaks and prompt injection “with attack success rate above 90% for most”; most had originally reported near-zero attack success. For agents, OWASP calls the least-privilege and capability-budgeting controls “load-bearing”.

Where detection and monitoring fit

Input filters lower the odds; they do not close the door. OWASP notes that rephrasing or encoding evades semantic filters, a common part of AI guardrails. Monitoring does a different job: the NCSC advises logging “potentially including full input and output of the LLM, as well as tool use, API calls”. An attacker will likely have to hone an attack, so failed tool calls could reveal it early. The NCSC’s buying advice: “Beware any that claim they can ‘stop’ prompt injection”.

Questions and answers

Can prompt injection be prevented?

Not reliably. OWASP states that no reliable prevention mechanism for prompt injection exists today, and the UK's NCSC argues it may never be mitigated the way SQL injection is. The defences that hold limit what a fooled model can do rather than stop it being fooled.

Does a prompt injection classifier stop attacks?

No classifier stops prompt injection reliably. Nasr et al. bypassed 12 recent defences against jailbreaks and prompt injection, several of them detectors, most with attack success above 90 percent, using attackers who adapted to each defence. OWASP notes that rephrasing or encoding evades semantic filters, and the NCSC advises wariness of any product that claims to stop prompt injection.

Sources

  1. LLM01:2026 Prompt InjectionOWASP GenAI Security Project, 4 Aug 2026
  2. Prompt injection is not SQL injection (it may be worse)UK National Cyber Security Centre, 8 Dec 2025
  3. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses against LLM Jailbreaks and Prompt InjectionsarXiv, 10 Oct 2025

Guides that use this term