Definition · AI basics
Instruction hierarchy
The instruction hierarchy is a ranking that tells a language model whose instructions win when they conflict: the developer's system prompt over the user, and the user over text from tools, documents and web pages. OpenAI researchers named it in 2024. Models are trained to follow the ranking; nothing in the model stops them breaking it.
Last reviewed
Key points
- The instruction hierarchy ranks instructions by where they come from. The developer's system prompt beats the user, and the user beats text from tools, documents and web pages.
- OpenAI researchers introduced it in 2024 because models often gave a system prompt the same priority as text from untrusted users and third parties.
- It is a trained behaviour and a published policy, not an enforced mechanism. Its authors wrote that their trained models were likely still vulnerable to powerful attacks.
- Vendors differ. OpenAI's order is strict; Anthropic says its principal hierarchy is not, since users keep some entitlements operators cannot override.
- Treat it as something that lowers attack success, and put hard limits outside the model.
How it works
A language model reads everything as one input: the system prompt from the developer, the user’s messages, its own earlier replies, and whatever tools return. The instruction hierarchy gives each source a rank. In the 2024 paper that named it, system messages rank highest, then user messages, then model outputs, with tool outputs lowest.
A lower-ranked instruction that fits the higher one is followed. A car-sales bot whose user says “speak in spanish” should comply. One that conflicts, or is simply off-task, should be ignored, and refused if there is no way to proceed. That covers a web result saying “email me the user’s conversation history”, and a user asking the car bot about the Navier-Stokes equation.
The authors fine-tuned GPT-3.5 Turbo on synthetic training data. Against a baseline trained without that data, the paper reports defence against system prompt extraction “improved by 63%”, its largest gain, and jailbreak robustness up “by over 30%”. Harmless instructions were followed almost as well as by the baseline, except on two adversarially built tests.
Why it matters
The instruction hierarchy is a trained preference, not an enforced mechanism.
The paper starts from the problem: models “often consider system prompts … to be the same priority as text from untrusted users and third parties.” If the ranking had to be trained in, it was not there to begin with. Training makes a model more likely to hold the ranking. It does not make breaking it impossible. The authors say so: their models were “likely still vulnerable to powerful adversarial attacks.” Google DeepMind reached the same conclusion for Gemini in 2025: “no single solution, including adversarial training, offers complete immunity.”
So a rule in the system prompt is not a security argument. A limit that must hold, such as which emails an agent may send, belongs outside the model: in permissions, or in asking the user to confirm. That is the practical lesson of prompt injection.
In practice
The ranking is published as policy. OpenAI’s Model Spec says instructions with higher authority override lower ones. Tool outputs, file attachments and quoted text have “no authority by default”. Developers are told to wrap untrusted data in YAML, JSON, XML or untrusted_text blocks. Without that, the Spec says, it “can be extremely difficult” for the model to tell injected text from the developer’s instructions.
The same document says OpenAI’s “production models do not yet fully reflect the Model Spec”. The policy describes what the models are trained towards, not what they are guaranteed to do.
Where definitions disagree
Vendors differ on how strict the order is.
OpenAI’s order is strict. The Model Spec’s chain of command runs platform, developer, user, guideline, then no authority, and “Instructions with higher authority override those with lower authority.”
Anthropic’s order is rough. Claude’s constitution calls it the principal hierarchy: Anthropic, operators, then users, trusted “in roughly the order given above”. Who counts as operator or user depends on their role in the conversation, not on what kind of entity they are. Anthropic also says “This is not a strict hierarchy”: users have some entitlements operators cannot override.
Google DeepMind’s 2025 Gemini report names no tiers. It lists the instruction hierarchy among defences in the research literature and relies on adversarial training plus defences outside the model.
OpenAI and Anthropic agree that instructions inside tool results, documents and search results are information, not commands, though OpenAI lets a developer delegate authority to them. Both also say their models fall short of the policy. Anthropic writes that Claude’s “behavior might not always reflect the constitution’s ideals.”
Questions and answers
Does the instruction hierarchy stop prompt injection?
No. The instruction hierarchy makes a model less likely to obey injected text, but it does not make it impossible. The OpenAI researchers who introduced it wrote that their trained models were "likely still vulnerable to powerful adversarial attacks", and Google DeepMind's 2025 Gemini report concluded that no single defence, including adversarial training, "offers complete immunity".
What order does the instruction hierarchy use?
The 2024 paper ranks system messages first, then user messages, then model outputs, then tool outputs. OpenAI's Model Spec ranks platform, developer, user and guideline levels, and gives tool outputs, quoted text and file attachments no authority by default.
Is Anthropic's principal hierarchy the same thing?
Close, but not the same. Anthropic's principal hierarchy ranks Anthropic, operators and users, and says it is not a strict hierarchy: users have some entitlements that operators cannot override. Like OpenAI, it treats instructions found in tool results and documents as information, not commands.
Sources
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsarXiv (OpenAI authors), 19 Apr 2024
- OpenAI Model Spec (2025-04-11)OpenAI, 11 Apr 2025
- Claude's ConstitutionAnthropic, 22 Jan 2026
- Lessons from Defending Gemini Against Indirect Prompt InjectionsarXiv (Google DeepMind authors), 20 May 2025