What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

AgentDojo

AgentDojo is an open-source framework for testing whether AI agents resist prompt injection. It gives an agent realistic tasks in simulated workspace, Slack, travel and banking environments, hides attacker instructions in the data the agent's tools return, and measures both how often the agent completes the user's task and how often it obeys the attacker.

Last reviewed

Key points

  • AgentDojo is an open-source framework from researchers at ETH Zurich and Invariant Labs, published in June 2024, for testing how AI agents hold up against prompt injection.
  • It gives an agent 97 realistic tasks across four simulated environments and hides attacks in the data its tools return, making 629 security test cases.
  • It scores two things: whether the agent finished the user's task, and whether it carried out the attacker's.
  • A score describes only the attacks that were tried. Against an agent built on a model AgentDojo rated among the most resistant, US and UK government red teamers raised attack success on held-out tasks from 11 percent for the strongest baseline attack to 81 percent for their strongest new one.

How AgentDojo works

AgentDojo builds four simulated environments: Workspace (email, calendar and a cloud drive), Slack, Travel and Banking. Each has tools an agent can call, such as reading a calendar or sending a payment, and a state filled with dummy data.

A test pairs two tasks. The user task is the job the agent was given, such as paying a bill. The injection task is the attacker’s goal, such as sending money to the attacker’s account. The attack text sits in data a tool returns, so the agent meets it mid-job, which is how indirect prompt injection works outside a lab. Pairing every user task with every injection task in an environment gives the 629 security test cases.

AgentDojo judges the outcome with fixed checks on the environment’s state and the agent’s answer, not by asking another model. Did the payment go to the right account? Did money leave for the attacker’s? That gives two scores: utility, and the targeted attack success rate.

The framework is meant to grow. Its authors built it as an extensible environment for new tasks, attacks and defences, not a fixed test suite.

Why AgentDojo matters

AgentDojo gives agent hijacking a shared yardstick. Two defences, or two models, can be compared on the same tasks and the same attacks, and the scoring is done by fixed checks rather than another model’s judgement.

The limit is the one in the analogy. A score covers the attacks in the test and no others. NIST puts it plainly: “A new system may be robust to attacks tested in previous evaluations, but real-life attackers can probe the new system’s unique weaknesses.”

What the measurements show

NIST’s Center for AI Standards and Innovation (CAISI) calls AgentDojo “a leading open-source framework for testing the vulnerability of AI agents”. In January 2025 the same team, then the US AI Safety Institute, reported that it had fixed bugs in AgentDojo’s default hijacking scenarios and added injection tasks for remote code execution, database exfiltration and automated phishing.

CAISI tested agents built on Anthropic’s upgraded Claude 3.5 Sonnet, released October 2024, a model AgentDojo had found to be one of the top performers at resisting hijacking. With red teamers from the UK AI Security Institute, it developed new attacks on some Workspace tasks and measured them on held-out ones. Attack success rose from 11 percent for the strongest baseline attack to 81 percent for the strongest new one. The model did not change; the attacks did. The new attacks also worked in the other three environments.

Retrying moved the number too. Across five injection tasks, the average attack success was 57 percent on one attempt and 80 percent when each attack was tried 25 times.

The five-task average mixes tasks whose harm is not comparable: a harmless email to a stranger counts the same as mailing the user’s five largest files to a stranger and demanding a ransom. NIST notes that a low rate on its data exfiltration task does not mean that scenario can be left unmitigated.

AgentDojo also scores defences. CaMeL, from researchers at Google, Google DeepMind and ETH Zurich, extracts the control and data flows from the user’s trusted request, so that, in the authors’ words, untrusted data “can never impact the program flow”. Security policies checked at each tool call limit where the data can go. With OpenAI’s o3 model at high reasoning effort, CaMeL solved 77 percent of AgentDojo tasks with provable security, against 84 percent for the same model undefended. The guarantee rests on the user’s own request being trusted, and the authors report that CaMeL is still open to side-channel attacks.

Questions and answers

What does AgentDojo measure?

AgentDojo measures two things for an AI agent: utility, whether it completes the user's task, and targeted attack success rate, how often an instruction hidden in tool output makes it carry out the attacker's goal.

Does a low AgentDojo attack success rate mean an agent is safe?

No. The rate describes only the attacks that were run. US and UK government red teamers tested an agent built on a model AgentDojo had found among the most resistant. On held-out tasks, attack success rose from 11 percent for the strongest baseline attack to 81 percent for the strongest new one.

Sources

  1. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsarXiv, 19 Jun 2024
  2. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsarXiv, 19 Jun 2024
  3. Technical Blog: Strengthening AI Agent Hijacking EvaluationsNIST Center for AI Standards and Innovation, 17 Jan 2025
  4. Defeating Prompt Injections by DesignarXiv, 24 Jun 2025