What matters in AI.

Subscribe

Learn / AI agents

Guide · AI agents

AI agent attacks

An AI agent reads untrusted text, decides what to do, and then acts with your credentials. Every named attack on agents targets one of those three steps, and the step it targets decides which defence works. These are the attacks, grouped that way.

Last reviewed

What the agent reads

Attacker text placed where the agent will read it and act on it, usually as instruction.

  • Agent hijacking: instructions hidden in data the agent ingests.
  • Structured template injection: forged chat-template turn markers in retrieved content, so attacker text reads as a user request.
  • BragJack: a browser extension takes over the page an in-browser agent trusts for orders, so attacker prompts arrive as the user’s.
  • Prompt Forcing: a lower-trust component writes the agent’s prompts itself, as if the user sent them, and can answer when the agent hesitates.
  • WARP: a persuasive passage appended to a community page research agents keep retrieving.
  • WebInject: a faint, optimised pixel change in a webpage that steers a screenshot-reading agent to one action.
  • Agent data injection: corrupts the trusted metadata the agent relies on.
  • Agent skill poisoning: a payload hidden in a skill file, trusted from install.
  • AI agent tool poisoning: a trusted tool’s description, code or output turned against the agent.
  • MCP tool poisoning: instructions in an MCP tool’s description that the approver never sees.
  • MCP tool shadowing: a rogue tool named close enough to a trusted one to be picked.
  • ToolHijacker: a new tool whose description is optimised so the agent’s retriever and model pick it for one task.
  • LLM Scope Violation: untrusted content pulls privileged data into the model’s context.
  • LLM prompt self-replication: an injected prompt the model copies into its own output, so it spreads to whatever reads that output.
  • ObliInjection: an injection optimised to work wherever the attacker’s segment lands among the sources assembled into the agent’s input.

What the agent decides

The input may be clean at the moment of attack; what has been tampered with is memory, planning, or the agent’s sense of when it is done.

What the agent can reach

Blast radius. Once the agent is redirected, the damage depends on what it was allowed to touch and who trusts it downstream.

Terms in this guide

In the news