Definition · AI security
Role confusion
Role confusion is a failure in how a language model identifies who is speaking. The model infers a message's role from how the text sounds, such as its style, wording or a claim about its source, rather than from the role tag the system assigned, so untrusted text that imitates a trusted role inherits that role's authority.
Last reviewed
Key points
- Role confusion was named in a February 2026 paper that traces prompt injection to it. Untrusted text which imitates a trusted role inherits that role's authority.
- Models read style, not tags. Probes trained only on role tags reported injected text as the role it imitated, and rephrasing an attack out of that style dropped its success from 61% to 10%.
- Two named attack families, on the narrow reading. CoT Forgery imitates the model's own reasoning; chat-template attacks such as ChatInject imitate the syntax that marks a user message.
- Forging a marker does not always beat asking plainly. Against six coding-agent harnesses in August 2026 both families scored 0 of 13 on five of them, and were often easier to spot.
- A second 2026 preprint uses the name much more narrowly, as one attack class beside prompt injection rather than the mechanism beneath it. Check which scope a source means.
A language model gets role tags — system, user, tool — recording where each piece of text came from. When a tag disagrees with how the text sounds, the sound wins. That gap is role confusion, and an attacker can write in any voice.
How a model decides who is speaking
Role tags encode privilege. In the naming paper’s words, “System prompts outrank
user messages; tool outputs should inform but not command” — a tool result is
“outside content which should be incorporated, but never obeyed”.
That paper measured whether models read the tags at all. It trained probes on
identical text wrapped in different tags, so a probe could only learn what a tag
does inside the model. Pointed at prompt-injected text, the probes reported the
role the text imitated, not the one it was tagged with. Put the model’s own
reasoning inside user tags and it still registers as reasoning.
Style and tag land in the same place inside the model; under conflict, style wins.
Why it matters
Role confusion makes the role tag unenforceable, so every defence resting on one — the instruction hierarchy, an agent’s rule that tool output is data — is declared at the interface and not enforced inside the model.
It also changes what prompt injection is: not a model failing to separate instructions from data, but a model separating them by a feature the attacker writes. The probes measure it before generation: across 1,000 agent exfiltration attempts, success rose from 2% in the least-confused band to 70% in the most.
The obvious defence costs the role its point. Train a model to spot forged reasoning and, unable to confine that suspicion, it must distrust all reasoning-like text, its own included.
In practice
On the narrow reading, two families forge two different markers. The wide reading does not stop at two: the naming paper finds that style, “while dominant, is not the only signal the model uses to infer role”, and that a cruder one works too — “you can simply tell the model what role you want in plain text, and it believes you”. It generalises past any list of attacks: “any forgeable signal becomes a surface for attacks or behavioral failures”.
CoT Forgery writes fake reasoning into a user message or a tool result, and the model adopts it as its own chain of thought. It is the naming paper’s demonstration attack, run zero-shot against six models with near-perfect safety benchmark scores. On StrongREJECT’s 313 harmful requests, raw prompts and standard jailbreaks achieved 0–4%; the forgery took gpt-oss and o4-mini past 80%, and the hardened GPT-5 family to between 17% and 52%. Against an agent with shell access, ordinary injection got 0–2% — except on gpt-oss-20b, where it got 26% — and the forgery 56–70%.
The forged reasoning does not have to make sense. Absurd justifications — that cocaine synthesis is fine because the user is wearing a green shirt — worked as often as plausible ones, 60% against 63%. What it has to do is sound right: rephrasing the same argument out of the model’s reasoning style dropped success from 61% to 10%, and changing the single phrase “The user” cost 19 points.
Structured template injection forges the other marker, writing the payload in the chat template’s own turn syntax so the model reads it as a user message. ChatInject, the named technique, raised average injection success on two agent benchmarks from 5% to 32% and from 15% to 46%, and kept working against closed-source models whose templates are not published. Phantom automates the search for effective templates and reports over 70 vendor-confirmed vulnerabilities in shipping products.
Forging a marker is not reliably better than asking plainly. Both families were
re-tested in August 2026 against six coding agents at named versions, with the
payload planted in a README, thirteen objectives and ten attempts each. Both
achieved nothing on five of the six harnesses, and nothing on any of the three
that offer automatic permission review. That paper’s reading is that the forgery
is conspicuous: every model it tested recognised the fake reasoning as an
injection, and Kimi named the fake <|im_user|> markup as the attack. Two
caveats belong with that. Those were reimplementations rather than the original
authors’ runs, and a coding agent reading a README is a narrower target than a
browsing agent or a chat model.
Where definitions disagree
The name is used at two scopes, and the difference is not cosmetic.
Prompt Injection as Role Confusion (ICML 2026) uses it for the model-side defect: how a model decides who is speaking, and why that decision is forgeable. Under that reading prompt injection is a consequence of role confusion, and so are failures that are not injection at all — the paper shows the same probes explaining why a long system prompt loses authority over the tokens at its end.
When Context Gets Root, an August 2026 preprint on agent harnesses, uses “role-confusion attacks” for a specific class: attacks that forge a role marker, listed beside prompt injection rather than above it. That narrow sense is the one in which role confusion contrasts cleanly with instruction privilege escalation, where nothing is forged because the harness genuinely hands the content a higher role.
Neither is wrong; they are answering different questions. One asks what is broken in the model, the other asks how to sort attacks on a harness. A claim about “role confusion” is worth checking against the scope its source meant.
Questions and answers
Is role confusion the same as prompt injection?
Not the same, but the two are closely tied, and how closely depends on which source you read. The February 2026 paper that named role confusion argues that prompt injection is a consequence of it, its title being "Prompt Injection as Role Confusion". On that reading, role confusion is the model-side defect and prompt injection is what an attacker does with it. An August 2026 preprint on agent harnesses uses the name more narrowly, for attacks that forge a role marker, and sets those beside prompt injection as a separate class. Neither is settled usage, and the narrow one is not even uniform within agent security: Phantom, a February 2026 paper on the same kind of attack, uses the name for the effect its forged templates induce rather than for a class they belong to. What can be said is that the wide reading is the one with the measurements behind it.
Does an instruction hierarchy stop role confusion?
No, because an instruction hierarchy is built on the thing role confusion breaks. A hierarchy assigns privilege by role tag, so it only constrains behaviour if the model reads the tag. The 2026 paper that named role confusion trained probes on nothing but tags and still found them reporting injected text as the role it imitated, and found that removing every tag from a conversation left the model's internal picture of who was speaking almost unchanged. The hierarchy is declared at the interface and not enforced inside the model.
Can you fix role confusion by making the role tags more distinctive?
No source claims that works, and the paper that named role confusion stops short of proposing any fix. It states the requirement rather than a mechanism, that role perception is governed by attacker-controllable features and that this is what a defence has to solve. It also argues that the obvious alternative, training the model to recognise forged text by its patterns, costs the role its purpose, since a model that cannot place a boundary around its suspicion must distrust all reasoning-like text, including its own.
Which attacks are role confusion attacks?
Two named families, by the narrow reading. CoT Forgery injects fabricated reasoning that the model adopts as its own chain of thought. Chat-template attacks write the payload in the model's own turn syntax so it reads as a user message, with ChatInject the named technique and Phantom an automated search for the templates that work best. Both forge a marker of a more privileged speaker. By the wide reading, ordinary indirect prompt injection also qualifies, because a command in a web page succeeds by sounding like a user rather than like retrieved data.
Sources
- Prompt Injection as Role ConfusionarXiv, 22 Feb 2026
- Prompt Injection as Role ConfusionarXiv, 22 Feb 2026
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsarXiv, 26 Sep 2025
- Automating Agent Hijacking via Structural Template InjectionarXiv, 18 Feb 2026
- When Context Gets Root: Privilege Escalation in LLM HarnessesarXiv, 27 Aug 2026