What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

CaMeL

CaMeL (CApabilities for MachinE Learning) is a prompt injection defense for AI agents from researchers at Google, Google DeepMind and ETH Zurich. One model writes the user's request as a program without seeing untrusted data, and an interpreter tags each value with its source and allowed readers, blocking tool calls that break a security policy.

Last reviewed

Key points

  • CaMeL turns the user's request into a program before the agent reads any untrusted data, so text hidden in an email or web page cannot change which steps run.
  • Every value carries tags for where it came from and who may read it, and a policy checked before each tool call stops a flow it forbids.
  • On the AgentDojo benchmark, at most 11 of 949 attacks per model tested still succeeded against CaMeL, and the authors place all of them outside CaMeL's threat model.
  • The protection costs completed work. With no attack running, OpenAI's o3 completed 77 percent of tasks under CaMeL against 84 percent undefended; Gemini 2.5 Pro fell from 73 to 41 percent.
  • CaMeL does not aim to stop injected text that only misleads the reader, such as a false summary or a phishing link. Its authors warn that side channels may still leak private data, and say prompt injection is not solved.

How CaMeL works

CaMeL splits an AI agent into parts that do not share what they read.

A privileged model sees only the user’s request. It writes the request as a short Python program: find the meeting notes, pull out Bob’s address, send him the document. It never sees what the tools return.

A quarantined model reads the untrusted data, such as the meeting notes, and fills in the fields the program asked for. It has no tools, and it cannot send messages back to the privileged model, because such a message could carry an injection.

A custom interpreter runs the program. It records which values each variable came from and tags every value with capabilities: its source and who may read it. Before each tool call, it checks the arguments against security policies. A confidential file headed for an address taken from someone else’s document fails the check.

The plan is written before any untrusted data is read, and every tool call passes a policy check.

Why CaMeL matters

Many prompt injection defenses try to make the model ignore injected instructions: delimiters, reminders of the task, extra training. The authors note that none of these give “any guarantee of security”. CaMeL puts the defense outside the model. The policy check runs on every tool call whether or not the model was fooled.

Isolation alone is not enough. In the Dual LLM pattern that CaMeL builds on, an injection in the notes can still make the quarantined model return the attacker’s address, so the file leaks. The policies close that gap.

On AgentDojo, at most 11 of 949 attacks per model tested still succeeded against CaMeL. The authors place all of them outside CaMeL’s threat model, such as an injected hotel review simply shown to the user.

What CaMeL does not stop

CaMeL’s authors list attacks it does not aim to stop. Injected text that only changes what the user reads is out of scope: an email summarized as something it does not say, or a planted message urging the user to click a malicious link so as not to lose their account. CaMeL also trusts the user’s own request.

The authors also find CaMeL “vulnerable to side-channel attacks”: ways to infer private data from how the system behaves, such as whether a program finishes. Their build blocks or mitigates the examples they show, but they believe other kinds might be possible and warn that data “can still be (inefficiently) mishandled”. Their own verdict on whether prompt injection is solved: “No”.

The released code is not a product. Its README warns the interpreter “likely contains bugs” and “might not be fully secure”, and says the authors are not planning to maintain it.

Trade-offs

Fewer tasks get done. With OpenAI’s o3 at high reasoning effort and no attack running, CaMeL completed 77 percent of AgentDojo tasks against 84 percent undefended. The authors’ abstract calls this “provable security”. The cost varies by model: Claude 4 Sonnet fell from 87 to 74 percent, and Gemini 2.5 Pro from 73 to 41 percent. No model tested could complete a request to “do the actions specified in the email”, because the privileged model never reads the email.

People must approve blocked actions. In a real deployment a call the policy forbids asks the user to confirm. The authors warn that users can grow numb to such prompts and approve malicious actions, and that someone has to write and maintain the policies. See human in the loop.

It costs more. With Claude 3.5 Sonnet, CaMeL used roughly two to three times the tokens of a plain tool-calling agent on the median task, with or without an attack running.

Questions and answers

Does CaMeL stop all prompt injection attacks?

No. CaMeL's authors answer that question themselves: prompt injection is not fully solved. CaMeL does not aim to stop injected text that only misleads the user, such as a false summary or a phishing link, and the authors show side-channel attacks that can still leak private data.

How is CaMeL different from the Dual LLM pattern?

CaMeL builds on the Dual LLM pattern, where a privileged model plans and a quarantined model with no tools reads untrusted data. CaMeL adds tags on every value and security policies at each tool call, because injected data can still change a tool's arguments, such as the recipient of an email, even when the plan stays fixed.

Can I deploy CaMeL in production?

Not as released. The CaMeL code on GitHub is a research artifact for reproducing the paper. Its authors warn the interpreter likely has bugs and might not be fully secure, and say they are not planning to maintain it.

Sources

  1. Defeating Prompt Injections by DesignarXiv, 24 Jun 2025
  2. CaMeL: Defeating Prompt Injections by Design (README)Google Research on GitHub