Definition · AI security
Prompt injection
Prompt injection is an attack on language models in which instructions are hidden inside text the model reads, causing it to follow the attacker rather than the operator. It works because a model receives instructions and data through the same channel and has no reliable way to tell them apart.
Last reviewed
Key points
- Prompt injection is the first entry on OWASP's Top 10 for LLM applications, and has been in every edition of that list.
- It works because a language model cannot tell instructions from data. OWASP says both are tokens on the same stream, so there is no clean equivalent to a parameterized query.
- Better prompting does not fix it. An instruction telling the model to ignore injected instructions is text in the same channel as the attack, and can be overridden the same way.
- Detection is a signal, not a fix. Twelve recent defences were re-tested against attackers who had read them, and most fell with attack success above 90 percent despite originally reporting near zero.
- The dangerous form is indirect: the instruction sits in a page, a document or a tool result the model retrieves, so the attacker never has to breach the application and the user never sees the instruction. The defences that hold limit what a bypassed model can do rather than stop it being fooled.
Prompt injection hides instructions inside text a language model reads. Whoever controls that text can give the model orders, because the model cannot tell it apart from the operator’s own instructions. Simon Willison named the attack in September 2022.
Direct injection is a user typing instructions that contradict the system prompt. Indirect prompt injection is the dangerous case: the instruction sits in content the model retrieves, so the attacker never has to breach the application and the user never sees the instruction.
Why prompt injection has no clean fix
Prompt injection has no clean fix because the boundary it violates was never there. OWASP states that models make “no architectural distinction between ‘instructions’ and ‘data’”, both being “tokens on the same stream”, so “there is no clean equivalent to parameterized queries.” Nothing in that stream carries a mark the model must respect.
Better prompting therefore cannot close it: an instruction to ignore injected instructions is text in the same channel as the attack. Nor can detection — twelve recent defences, re-tested against attackers who had read them, mostly fell with attack success above 90 percent.
The NCSC frames a language model as a confused deputy that is inherently confusable, so the risk can be reduced and not removed. OWASP agrees: no reliable prevention mechanism exists, so defence is “architectural rather than interceptive”, with least privilege “load-bearing” for agents. What holds is prompt injection mitigations, and for an agent, the lethal trifecta.
Where definitions disagree
Whether jailbreaking is prompt injection is disputed. OWASP says it is, in the 2026 edition and the 2025 one before it: jailbreaking is “the subset of prompt injection where the attacker’s goal is to make the model violate its safety protocols.” Willison, who coined the term, says it is not: with no concatenation of a trusted developer prompt and untrusted input, it is not prompt injection.
This page follows the narrow reading, because the parties and the defences differ. Jailbreaking is the operator pushing on the model’s own safety training; injection is a third party pushing on someone else’s application. Willison’s example of the cost: a detector trained on jailbreaks blocks a request for a napalm recipe and allows an instruction to search a mailbox and forward the results to an attacker.
Questions and answers
Is prompt injection the same as jailbreaking?
Authorities disagree. OWASP's LLM Top 10 treats jailbreaking as the subset of prompt injection where the attacker's goal is to make the model violate its safety protocols. Simon Willison, who coined the term, keeps the two apart, and this page follows him: jailbreaking attacks the safety training inside the model, usually driven by the person operating it, while prompt injection attacks an application that concatenates a trusted prompt with untrusted text, and the attacker is usually someone other than the operator.
Can prompt injection be fixed by better prompting?
No. Instructing a model to ignore injected instructions is itself text in the same channel the attack uses, so it can be overridden the same way. OWASP classes constraining the model in the system prompt as a partial control only, to be paired with privilege controls held in application code. The defences that hold limit what the model is able to do rather than stop the model being fooled.
Is prompt injection the same as SQL injection?
Prompt injection and SQL injection share a shape and not a cure. Simon Willison named prompt injection after SQL injection because both concatenate untrusted input with trusted instructions. The UK's National Cyber Security Centre argues the comparison becomes dangerous past that point: parameterised queries fix SQL injection at its root because a database engine enforces the line between instruction and data, and a language model enforces no such line, so the mitigation has no equivalent.
Sources
- Prompt injection attacks against GPT-3Simon Willison, 12 Sep 2022
- LLM01:2026 Prompt InjectionOWASP GenAI Security Project
- LLM01:2025 Prompt InjectionOWASP GenAI Security Project
- OWASP Top 10 for LLMs and Gen AI Apps 2023-24OWASP GenAI Security Project
- Prompt injection and jailbreaking are not the same thingSimon Willison, 5 Mar 2024
- Prompt injection is not SQL injection (it may be worse)UK National Cyber Security Centre, 8 Dec 2025
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsarXiv, 10 Oct 2025