Definition · AI agents
Dual LLM pattern
The Dual LLM pattern is a design for AI assistants that resists prompt injection by splitting the work between two language models. A privileged model takes the user's request and calls tools but never sees untrusted content. A quarantined model reads that content but has no tools, and its output is stored in variables the privileged model handles by name.
Last reviewed
Key points
- The Dual LLM pattern splits an AI assistant in two: a privileged model that has tools but never reads untrusted content, and a quarantined model that reads untrusted content but has no tools.
- Ordinary software, the controller, stores what the quarantined model returns in variables, so the privileged model handles names like $VAR1 rather than the text itself.
- The pattern keeps injected text away from the model that plans and calls tools, but not from the data the plan uses. If the quarantined model is fooled into extracting an attacker's address, a planned send step sends the file there.
- Simon Willison proposed the pattern in April 2023 and called it "pretty bad" himself. The authors of CaMeL (2025) describe their system as, to the best of their knowledge, its first concrete instantiation.
How the Dual LLM pattern works
The Dual LLM pattern divides an AI agent into three parts.
The privileged LLM takes input only from trusted sources, mainly the user. It has the tools, such as sending email.
The quarantined LLM handles anything untrusted, such as an email or a web page. It has no tools. Simon Willison, who proposed the pattern in April 2023, says it “is expected to have the potential to go rogue at any moment”.
The controller is ordinary software, not a language model. It runs the tools, calls the quarantined LLM and stores every result in a variable.
The quarantined LLM’s output does not reach the privileged LLM as free text. Asked to summarize the latest email, the privileged LLM has the controller fetch it into $VAR1, have the quarantined LLM summarize $VAR1 into $VAR2, and show the user $VAR2. Checked output, such as a label from a fixed list, is the exception.
The privileged LLM plans with variable names and never sees the email or its summary.
Why the Dual LLM pattern matters
The Dual LLM pattern keeps prompt injection away from the part of an assistant that can act.
The pattern protects the plan, not the data it runs on. In an attack from the CaMeL paper, planted text in meeting notes makes the quarantined model return an attacker’s file name and address. The plan runs as written, and the document goes to the attacker. In the authors’ words, “while the control flow is protected by the Dual LLM pattern, the data flow can still be manipulated”.
When a plan picks its next step from untrusted data, the data becomes the plan. Asked to follow an email’s instructions, models in the CaMeL tests wrote code that runs whatever the email says, which isolation alone cannot stop. CaMeL adds tags and security policies to narrow these gaps, without claiming a complete fix.
Trade-offs
Some tasks do not fit. When the actions depend on untrusted data, such as doing what an email asks, the privileged LLM cannot plan them safely, because it never reads the email. The CaMeL authors call this failure inherent to the Dual LLM pattern. In the authors’ tests, models tried workarounds and none solved that task.
The quarantined model can still leak what it reads. It has no tools, but it can still read any sensitive data it is given and write attacker-controlled output. For an email assistant whose quarantined model drafts replies, Beurer-Kellner and colleagues warn the draft could be tampered to encode sensitive data or carry instructions to the user’s coworkers.
The user is still a way in. The pattern trusts whatever the user types. A user tricked into pasting in untrusted content defeats it, and output shown to the user can itself talk them into copying their own data out.
It is hard to build correctly. Willison calls his own proposal “pretty bad”: more complex to build and worse to use. If new features cannot be added without mistakes that leak untrusted text to the privileged LLM, he warns, everything built for protection is wasted.
Questions and answers
Does the Dual LLM pattern stop prompt injection?
Not completely. The Dual LLM pattern stops injected text from reaching the model that can call tools. The quarantined model can still be fooled into returning attacker-chosen data, such as a different email address, which the planned actions then use. The pattern also trusts whatever the user types, so a user tricked into pasting in an attack defeats it.
What is the difference between the Dual LLM pattern and CaMeL?
The Dual LLM pattern is Simon Willison's 2023 design: a privileged model with tools, a quarantined model without them, and software passing variables between them. CaMeL, from Google, Google DeepMind and ETH Zurich, builds that design and adds tags on every value and security policies at each tool call, because the pattern alone leaves tool arguments open to manipulation.
Why can the privileged LLM not just read the quarantined LLM's output?
Because the output can carry the injection with it. A quarantined model that read a malicious email can repeat or rephrase the attacker's instructions in its answer. Willison allows one exception: output that can be checked, such as one label from a fixed list, may be validated and passed on.
Sources
- The Dual LLM pattern for building AI assistants that can resist prompt injectionSimon Willison's Weblog, 25 Apr 2023
- Defeating Prompt Injections by DesignarXiv, 24 Jun 2025
- Design Patterns for Securing LLM Agents against Prompt InjectionsarXiv, 27 Jun 2025