Definition · AI agents
MCP tool poisoning
MCP tool poisoning is an attack in which a Model Context Protocol server hides instructions inside a tool's description or its responses, so an agent that trusts the tool follows the attacker's commands. The model reads the hidden text; the person who approved the tool never sees it.
Last reviewed
Key points
- The trust gap is timing, not absence: a tool's description is reviewed once, at connection. Every later response it returns skips that check entirely.
- Invariant Labs' 2025 proof of concept hid an instruction inside a calculator tool, telling Cursor's model to read the user's SSH key and return it disguised as a calculation note; the person saw only a plain summary.
- This works by design. The model reads a tool's full text, hidden instructions included, while the editor shows the person only a simplified version.
- MCPTox tested 45 real MCP servers against 20 models and measured an average 36.5% attack success rate, reaching 72.8% for the most vulnerable model, with almost no refusals.
- It is one case of the wider AI agent tool poisoning technique MITRE ATLAS catalogues as AML.T0110, narrowed to attacks carried specifically over the Model Context Protocol.
How it works
A MCP server lists each tool with a short natural-language description meant to tell the connecting agent what the tool does. The model context protocol treats that text as data, but the model reads it in the same context as its instructions and can follow it, which is why OWASP classes the attack as a form of indirect prompt injection. Invariant Labs, who disclosed the technique in April 2025, hid an instruction inside a calculator tool’s description telling the connected model to read the user’s SSH private key and return it disguised as a calculation note. Tested against the Cursor editor, the model complied, while Cursor showed the person only a plain summary, never the hidden text.
The same gap sits on the response side, and it is harder to close: what a server returns on each call gets no review at all, so a server can turn honest into malicious at any later call.
Why it matters
MCPTox, an academic benchmark, tested 45 live MCP servers and 353 real tools against 20 LLM agents and found they followed a poisoned tool’s instructions 36.5% of the time on average, rising to 72.8% for the most vulnerable model. Agents almost never refused: even the most cautious model declined fewer than 3% of attempts. More capable models were often more susceptible, since the same instruction-following ability that makes a good agent also makes it a compliant one.
The impact scales with what the tool can already do: an agent that can read files or reach the network carries that access into every call, so a poisoned tool reaches as far as the agent’s own privileges do.
MITRE ATLAS catalogues the general technique as AML.T0110, AI agent tool poisoning, and treats the Model Context Protocol case as one instance of it: MCP made connecting an unreviewed tool server routine rather than exceptional.
Where definitions disagree
OWASP’s MCP Top 10 uses the same name, MCP03:2025 Tool Poisoning, but defines it differently: an attacker tampering with a tool’s schema or contract so a benign-looking operation maps to a destructive one, a supply-chain-style compromise of the interface itself. OWASP’s own community attack wiki defines tool poisoning the way Invariant Labs first described it: hidden instructions in a tool’s description or response text. The Top 10 entry lists such hidden instructions only as a detection signal. This page follows the original disclosure and the community wiki, the sense MCPTox also uses.
Questions and answers
Is MCP tool poisoning the same as prompt injection?
It is a specific channel for indirect prompt injection. The hidden instruction arrives through a tool's description or response instead of through user-supplied content or a web page.
Does approving an MCP server once protect against this?
No. Approval reviews the tool's description at connect time. What the tool actually returns on later calls is not checked again.
Sources
- MCP Security Notification: Tool Poisoning AttacksInvariant Labs, 1 Apr 2025
- MCP Tool PoisoningOWASP Foundation
- MCP03:2025 - Tool PoisoningOWASP Foundation
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP ServersarXiv (Wang et al.), 19 Aug 2025