What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

AI agent tool poisoning

AI agent tool poisoning is an attack that corrupts a tool an AI agent calls — its stated definition, its underlying implementation, or the responses it returns at runtime — so a call that looks routine carries out the attacker's instructions instead of, or alongside, its intended function. MITRE ATLAS classifies it as AML.T0110, a persistence technique.

Last reviewed

Key points

  • MITRE ATLAS names three ways a tool can be poisoned: its stated definition or instructions, its underlying implementation, or the responses it returns at runtime.
  • The attack lands on the tool, not the agent's prompt. Once a tool is added, the agent has no reason to re-examine it, so a later change to any of the three goes unnoticed.
  • MCP made the target common. ATLAS notes agent tools have "exploded in popularity" across thousands of publicly listed MCP servers, an ecosystem where "safeguards are often not in place."
  • It has already hit a shipping product. Security firm Koi Security identified a malicious postmark-mcp npm package in September 2025 that secretly copied every email it sent to an external address, which ATLAS now cites as its own reference case.
  • ATLAS maps no dedicated mitigation to this technique in the current collection, unlike several neighboring ones.

How it works

MITRE ATLAS names three places a tool an AI agent already trusts can be poisoned: its stated definition or instructions, its underlying implementation, or the response it returns at runtime. None of the three touches the agent’s own prompt or model weights — the attack lands entirely on the tool’s side of the call, which an agent has no routine reason to re-examine once the tool has been added.

OWASP describes the resulting gap for the common case, a tool served over the Model Context Protocol: a description is reviewed once, at connection time, while its responses “go straight into the LLM context with no equivalent check” on every later call. Invariant Labs, which first disclosed the pattern, adds a human-facing gap: a model sees a tool’s full description, hidden instructions included; a user usually sees only a simplified version. Indirect prompt injection is the usual channel that turns poisoned content into an action once it reaches the agent’s context.

Why it matters

Tool count has outgrown any process for re-checking them. ATLAS says agent tools have “exploded in popularity” across thousands of MCP servers, with “safeguards … often not in place.”

That is not hypothetical. In September 2025, Koi Security found what its CTO called “the world’s first sighting of a real-world malicious MCP server”: npm package postmark-mcp passed as a working tool through fifteen releases before one later version quietly BCC’d every email it sent to an outside address, downloaded 1,643 times before removal. Postmark, the real provider it impersonated, said it had never published such a package. ATLAS now cites the incident as its own reference case, an example of the AI supply chain compromise a poisoned tool travels through to reach a victim.

Questions and answers

Is AI agent tool poisoning the same as MCP tool poisoning?

MCP tool poisoning is the specific case where the poisoned tool is served over the Model Context Protocol. AI agent tool poisoning is the broader technique MITRE ATLAS defines as AML.T0110, covering any agent tool whose definition, implementation, or runtime responses have been corrupted, whatever protocol or integration delivers it to the agent.

How is this different from AI supply chain compromise?

AI supply chain compromise is the broader category of an adversary compromising how an AI artifact reaches a victim, covering models, data, and software generally. ATLAS lists AI Agent Tool as one specific compromise vector inside that category, one way a poisoned tool can be introduced. AI agent tool poisoning is the narrower technique that covers what happens once an agent actually invokes that tool.

Has AI agent tool poisoning happened outside a lab?

Yes. In September 2025, security firm Koi Security identified a package on npm named postmark-mcp that had passed as a working tool through 15 releases before a later version quietly added code that copied every email the tool sent to an outside address. Postmark, the real email service the package impersonated, said it had never published such a package. MITRE ATLAS now cites this incident as its own reference case for the technique.

Does anything mitigate AI agent tool poisoning?

MITRE ATLAS maps no dedicated mitigation to AML.T0110 in the current collection, unlike some neighboring techniques. Until one exists, the practical defense is the same one that applies to any supply-chain risk: checking where a tool comes from and who maintains it, not only what it appeared to do the one time it was reviewed.

Sources

  1. MITRE ATLAS, technique AML.T0110 AI Agent Tool Poisoning (collection 2026.08)MITRE
  2. MCP Tool PoisoningOWASP Foundation
  3. MCP Security Notification: Tool Poisoning AttacksInvariant Labs, 1 Apr 2025
  4. Security Alert: Malicious 'postmark-mcp' npm Package Impersonating PostmarkPostmark, 25 Sep 2025
  5. Malicious MCP server on npm postmark-mcp harvests emailsSnyk
  6. First Malicious MCP Server Found Stealing Emails in Rogue Postmark-MCP PackageThe Hacker News, 30 Sep 2025

Guides that use this term