Definition · AI agents
Agent hijacking
Agent hijacking is an attack in which instructions hidden in data an AI agent reads redirect the agent to act on the attacker's behalf. Agent hijacking is indirect prompt injection aimed at a system that can call tools, so the result is not a wrong answer but an unauthorised action taken in the operator's name.
Last reviewed
Key points
- Agent hijacking is indirect prompt injection against a system that can act. The attacker plants an instruction in data the agent reads, and the agent carries it out with the operator's credentials.
- What changes is the blast radius, not the mechanism. OWASP names agentic execution as the property under which the blast radius extends to whatever the agent's tools can reach.
- There is no user-visible step. In the EchoLeak pattern the victim never clicks anything, because the instruction arrives inside content the agent fetches on its own.
- Every frontier model tested has been hijacked. Across a public red-teaming competition against 13 frontier models, at least one successful attack was found against all of them.
- NIST scopes the term narrowly, as indirect injection against an agent. OWASP's ASI01 Agent Goal Hijack is wider and also covers a user talking an agent off its task interactively.
An agent hijack ends in a tool call. Somewhere upstream — an email, a web page, a file the agent fetched — an attacker left an instruction, and the agent carried it out with the operator’s credentials. NIST defines agent hijacking as a type of indirect prompt injection in which an attacker inserts malicious instructions into data that may be ingested by an AI agent, causing it to take unintended, harmful actions.
The term is new; the vulnerability is not. NIST calls it the latest incarnation of an age-old computer security problem, one that arises when a system lacks clear separation between trusted internal instructions and untrusted external data.
Why an agent is worse to hijack than a chatbot
The agent has tools. What changes is the blast radius, not the mechanism: OWASP names agentic execution as the deployment property under which the blast radius extends from the chat surface to whatever the agent’s tools can reach. A hijacked chatbot writes a bad summary. A hijacked agent sends mail, moves money or opens a pull request.
There is nothing for the victim to notice. Microsoft records CVE-2025-32711, known as EchoLeak, as an AI command injection in Microsoft 365 Copilot allowing an unauthorized attacker to disclose information over a network, scored 9.3 critical; published analysis describes a single crafted email that exfiltrated data with no user interaction.
Every frontier model tested has been hijacked. Analysing a public red-teaming competition against 13 frontier models, NIST reports that at least one successful attack was found against all of them. What helps is architectural rather than interceptive, which is OWASP’s position. Two framings of that architecture are the lethal trifecta and the Agents Rule of Two.
Where definitions disagree
How wide the term is decides what a claimed defence covers. NIST scopes agent hijacking narrowly, as indirect prompt injection against an agent. OWASP’s Top 10 for Agentic Applications ranks Agent Goal Hijack first, as ASI01, and scopes it wider: the attacker redirects the agent’s objectives or decision pathways “regardless of whether the manipulation occurs interactively or through pre-positioned inputs such as documents, templates, or external data sources”. A user who talks an agent off its task in conversation is ASI01, and is not NIST agent hijacking. The same document draws the far boundary, separating ASI01 from ASI06, memory and context poisoning, which is persistent corruption of stored context.
Questions and answers
Is agent hijacking the same as prompt injection?
Agent hijacking is a form of prompt injection, not a separate vulnerability. NIST defines agent hijacking as a type of indirect prompt injection in which an attacker inserts malicious instructions into data an AI agent ingests, causing it to take unintended, harmful actions. The distinguishing feature is not the attack but the target: an agent can call tools, so a successful injection produces an action rather than only an unwanted answer.
Can an agent be instructed not to obey instructions it finds in data?
No, not reliably. An instruction telling an agent to ignore commands found in retrieved content is itself text in the same context window as those commands, carrying no privilege the running model can enforce. OWASP's current entry, LLM01:2026, states that no reliable prevention mechanism exists today and that defence is therefore architectural rather than interceptive: design the surrounding system on the assumption that the instruction boundary will eventually be bypassed, and constrain what the agent is permitted to do and what its output is permitted to reach.
Has agent hijacking happened outside a laboratory?
Yes. CVE-2025-32711, published by Microsoft in June 2025 and known as EchoLeak, is recorded as an AI command injection in Microsoft 365 Copilot that allowed an unauthorized attacker to disclose information over a network, scored 9.3 critical. Published analysis describes a single crafted email that caused data exfiltration with no user interaction beyond normal use of the assistant.
Sources
- Technical Blog: Strengthening AI Agent Hijacking EvaluationsNIST Center for AI Standards and Innovation, 17 Jan 2025
- Insights into AI Agent Security from a Large-Scale Red-Teaming CompetitionNIST Center for AI Standards and Innovation, 23 Mar 2026
- CVE-2025-32711Microsoft, as CVE Numbering Authority, 11 Jun 2025
- EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM SystemarXiv, 6 Sep 2025
- OWASP Top 10 for Agentic ApplicationsOWASP Gen AI Security Project, 9 Dec 2025
- OWASP Top 10 for Agentic Applications 2026, ASI01: Agent Goal HijackOWASP Gen AI Security Project, 9 Dec 2025
- LLM01:2026 Prompt InjectionOWASP Gen AI Security Project
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationSimon Willison, 16 Jun 2025
- Agents Rule of Two: A Practical Approach to AI Agent SecurityMeta, 31 Oct 2025