Definition · AI agents
Instruction privilege escalation
Instruction privilege escalation is an attack on AI agent harnesses in which attacker-controlled content that entered at a low instruction level, such as a tool result, is later placed by the harness into a higher-privileged position, such as a user or system message, so the model follows it as authorised instruction.
Last reviewed
Key points
- Instruction privilege escalation was named in an August 2026 preprint, which defines it as context construction presenting content at a higher instruction level than the level it originally entered at.
- The failure is in the harness, not the model. When an agent delegates a task, the harness writes the delegated text into the subagent's context as a user message and drops where it came from.
- This is not prompt injection. No role marker is forged; the content arrives under a genuine user or system role that the harness itself assigned.
- Automatic permission review does not stop it. In the paper's example the reviewer rated the action high risk and allowed it anyway, because the transcript showed a user asking for it.
- Across six coding-agent harnesses the paper reports all 13 attack objectives achieved on every harness under full access, and on all three that offer automatic permission review.
A coding agent reads an untrusted file and treats it as untrusted. Then it hands a step to a subagent, and the harness writes that same text into the subagent’s context as a user message. The words did not change. Their privilege did.
How it works
Instruction privilege escalation exploits the moment a harness rebuilds context for a new model call. The 2026 preprint that named it maps the message types of four model APIs onto three levels — tool below user below system-effective — and defines the attack as context construction presenting content at a higher instruction level than its origin level.
Delegation is the clearest route. A main agent reads attacker text at the tool level, is induced to pass a task to a subagent, and the harness inserts that task as a user message. The same transition happens through persistent goals, scheduled tasks and invoked skills. A planted custom subagent goes further: its system prompt loads as system-effective policy, which the authors call tool-to-system escalation.
Why it matters
Two defences fail at once, and both behave correctly. The working agent obeys the visible hierarchy — it simply sees a user instruction. The permission reviewer reads the same transcript to infer who authorised an action. In the paper’s motivating case the reviewer classified starting a backdoored server as high risk and approved it anyway, on the grounds that “the user explicitly approved this exact action”.
This is why the paper separates the attack from prompt injection. In prompt injection the malicious content stays at its original level and argues its way past the model. In role confusion attacks it fakes a role marker. Here nothing is forged: the harness assigns the higher role, and the instruction hierarchy the model was trained to follow does the rest.
In practice
The evaluation covered 13 attack objectives — spanning confidentiality, integrity, availability and remote code execution — against six coding-agent harnesses: Claude Code, Codex, Gemini CLI, Qwen Code, Kimi and OpenCode, each at a named version. Under full access, escalation achieved all 13 objectives on every harness. Under automatic permission review it achieved all 13 on all three harnesses that offer that mode, where three published tool-level attacks — one prompt injection, two role confusion — achieved none.
Two limits belong with those numbers. Each case was “capped at 10 attempts”, so achieving an objective means achieving it within ten tries, not reliably; end-to-end success rates ranged from 31.7% to 100%, and installation of a malicious custom subagent succeeded in 46.2% to 100% of trials under automatic permission review, with Claude Code at the low end. The attacks were also handicapped on purpose. Every one had to start at the tool level, and the malicious code stayed plain-text Python with no obfuscation, so the model could see what it was being asked to run. It ran it anyway. In one case a Claude Code subagent deleted an application database and only afterwards asked whether that had been intended.
The paper proposes no fix. Its conclusion is that neither the models nor their permission reviewers prevent the attack.
Where definitions disagree
The phenomenon has two names and five days between them, and neither paper cites the other. When Context Gets Root, from Nanjing University and Honor, was posted on 27 August 2026 and calls it instruction privilege escalation. What’s in Your Agent’s Context?, from the University of Illinois Urbana-Champaign, was posted on 1 September 2026 and calls the same movement Message-Role Context Privilege Escalation, or M-CPE: attacker-controlled content from a low-privileged context ending up in a higher-privileged message role.
The scope is not identical. The Illinois paper adds a second axis the Nanjing one does not measure — Cross-Scope Context Privilege Escalation, where the content survives beyond the context it entered, for example by reaching a memory file that outlives the session. Under that taxonomy, a payload can escalate in role, in scope, or in both. The Nanjing paper tracks level only, and treats persistence as one mechanism among several rather than a separate dimension.
Neither name is settled. Both preprints are recent, unreviewed, and describe the same underlying defect: a harness that reassembles context without carrying the provenance of what it is reassembling.
Questions and answers
Is instruction privilege escalation just prompt injection with a new name?
No. The 2026 preprint that named instruction privilege escalation separates the two explicitly. In prompt injection, the malicious content stays at the instruction level it arrived at and tries to talk the model into following it. In instruction privilege escalation the content is moved to a higher level by the agent harness itself, so the model receives it under a real user or system role rather than as low-privilege text pretending to be one. The paper's baseline measurements make the practical difference visible. Three published tool-level attacks were tested first: one prompt injection and two role-confusion attacks, one of which wraps its payload in the model's own chat template. Under automatic permission review they achieved none of the 13 objectives on any of the three harnesses offering that mode, while escalation achieved all 13 on all three.
Does an instruction hierarchy prevent instruction privilege escalation?
Not on its own. An instruction hierarchy is a model-side defence that assigns privilege according to the source a message claims. The attack works because the harness, not the model, decides which level a piece of content is presented at when it builds the context for the next model call. Under escalation the model is not being fooled — the content genuinely arrives at user level — so obeying the hierarchy faithfully is what carries the attack. Once attacker text has been rebuilt into a user message, a model obeying the hierarchy is obeying the attacker.
Which agent features can cause this escalation?
The 2026 preprint names five and tests four. Delegating a task to a subagent inserts the delegated text into the subagent's context as a user message. Persistent goals and scheduled tasks re-deliver stored text as user-level prompts later. An installed custom subagent has its system prompt loaded as system-effective policy. Those four are the ones the evaluation covers. The fifth is an installed skill, which splits across both — metadata loaded as configuration, the body delivered as user-level instructions on invocation — but the paper describes that route from prior work rather than testing it. What these share is that each is a normal, documented feature being used as designed.