Guide · AI agents
AI agent attacks
An AI agent reads untrusted text, decides what to do, and then acts with your credentials. Every named attack on agents targets one of those three steps, and the step it targets decides which defence works. These are the attacks, grouped that way.
Last reviewed
What the agent reads
Attacker text placed where the agent will read it and act on it, usually as instruction.
- Agent hijacking: instructions hidden in data the agent ingests.
- Structured template injection: forged chat-template turn markers in retrieved content, so attacker text reads as a user request.
- BragJack: a browser extension takes over the page an in-browser agent trusts for orders, so attacker prompts arrive as the user’s.
- Prompt Forcing: a lower-trust component writes the agent’s prompts itself, as if the user sent them, and can answer when the agent hesitates.
- WARP: a persuasive passage appended to a community page research agents keep retrieving.
- WebInject: a faint, optimised pixel change in a webpage that steers a screenshot-reading agent to one action.
- Agent data injection: corrupts the trusted metadata the agent relies on.
- Agent skill poisoning: a payload hidden in a skill file, trusted from install.
- AI agent tool poisoning: a trusted tool’s description, code or output turned against the agent.
- MCP tool poisoning: instructions in an MCP tool’s description that the approver never sees.
- MCP tool shadowing: a rogue tool named close enough to a trusted one to be picked.
- ToolHijacker: a new tool whose description is optimised so the agent’s retriever and model pick it for one task.
- LLM Scope Violation: untrusted content pulls privileged data into the model’s context.
- LLM prompt self-replication: an injected prompt the model copies into its own output, so it spreads to whatever reads that output.
- ObliInjection: an injection optimised to work wherever the attacker’s segment lands among the sources assembled into the agent’s input.
What the agent decides
The input may be clean at the moment of attack; what has been tampered with is memory, planning, or the agent’s sense of when it is done.
- Agent memory poisoning: attacker content written into long-term memory.
- AI agent context poisoning: planted context that outlives the turn that introduced it.
- Delayed-trigger prompt injection: a dormant rule planted now that only activates when a later trigger phrase, event or condition appears.
- LoopTrap: corrupts the agent’s judgement of when a task is finished.
- STAC: harmless tool calls chained into a harmful outcome.
- Rogue agents: an agent whose behaviour has drifted from its purpose.
- Cascading failures: one fault amplified across agents and workflows.
What the agent can reach
Blast radius. Once the agent is redirected, the damage depends on what it was allowed to touch and who trusts it downstream.
- Excessive agency: an agent permitted to do more than its task needs.
- Tool misuse: a legitimate tool the agent is allowed to use, applied unsafely: deleting data, running up API costs or sending information out.
- Agentic identity and privilege abuse: an agent’s credentials, delegated tokens or inherited trust used beyond its task.
- Lethal trifecta: private data, untrusted content and a way out, held together.
- Workflow identity hijacking: a workflow acting with its own privileges, not the requester’s.
- Agent-to-agent lateral movement: one compromised agent reaching the others.
- Insecure inter-agent communication: messages between agents that are forged, altered or replayed, and trusted anyway.
- Malicious LLM proxy router: a third-party API router between agent and model provider that reads all traffic in plaintext, so it can rewrite tool calls before they run or copy the credentials passing through.
- Improper output handling: the agent’s output passed to a browser, database query or shell without validation or encoding, so whoever steers the agent reaches those systems.
- Unexpected code execution: agent code generation turned into code execution on the host.
- Human-agent trust exploitation: the person reading the output is the target.
Related topics
Terms in this guide
In the news
- Anthropic stops live internet access for its internal tests
- Some cyber insurance policies do not include AI agent attacks
- Wikimedia reports that OpenAI agents sent millions of API requests
- About 700 AI agents attacked Hugging Face, reports say
- Arbiter gives 152 results for the system prompts of 3 coding agents
- An OpenAI agent had unauthorised access to a Medicare portal
- Microsoft makes MXC available to limit the files an AI agent can use
- Incognia started a fraud check for AI agents
- Goodfire says its monitor for AI agents costs $51 for 1,500 sessions
- One chat message gave Zenity control of all AgentCore agents
- The first AWS patch for AgentCore agents did not remove the risk
- President Lee thinks AI agents were used in attacks on 7 or more firms
- A rule monitor found about 70% of attacks on AI agents in a test