Definition · AI agents
ToolHijacker
ToolHijacker is a prompt injection attack on how an LLM agent chooses its tools, introduced in a 2025 paper presented at NDSS 2026. The attacker publishes one crafted tool document whose description is optimised so the agent's retriever shortlists it and its language model then picks it for an attacker-chosen task.
Last reviewed
Key points
- ToolHijacker attacks tool selection, which tool an agent decides to use, not what a tool does once it runs.
- The attacker adds one new tool to the agent's library. Its description has two parts: one gets it shortlisted by the retriever, the other gets it picked by the model.
- The attacker never sees the target agent and tunes the tool on a stand-in pipeline instead. In the paper's headline test it still transferred: 96.7% of GPT-4o's picks for the target task on the MetaTool benchmark went to the malicious tool, and 88.2% on the larger ToolBench library.
- Models fine-tuned to resist prompt injection (StruQ, SecAlign) still picked the malicious tool 84.6% to 99.6% of the time. Injection detectors missed most malicious documents; the authors say the text reads as an ordinary on-topic description.
- All results come from two academic benchmarks. The paper reports no attack on a live tool marketplace.
How it works
Many AI agents choose a tool in two steps. A retriever compares the user’s request with every tool description in a library and shortlists the closest few. A language model then picks one tool from that shortlist.
ToolHijacker attacks both steps with one document. The attacker publishes a tool, say on an open hub, whose description has two parts. The first part matches the many ways a user might phrase the target task, so the retriever shortlists it. The second makes the model pick it from that shortlist.
The attacker has no access to the target agent. They build a stand-in (their own sample requests, retriever, model and tool library) and tune against it, in one of two ways. In the first, an LLM writes the first part from the sample requests, and the second is rewritten for up to 10 rounds to win the stand-in model’s choice. The second uses the stand-in models’ gradients.
Why it matters
In the paper’s main results, with an open Llama stand-in, the tuned tool won 96.7% of GPT-4o’s choices for the target task on the MetaTool benchmark, and 88.2% on ToolBench. Across that table’s eight target models and two benchmarks, success never fell below 74.3%, though weaker stand-ins scored as low as 34%. The best earlier attack compared, PoisonedRAG, reached 39.3% and 58.3% in the same GPT-4o tests.
Among 9,650 genuine tools, one malicious tool made the retriever’s top five for over 96% of target requests.
Prompt injection defences did little. Models fine-tuned to ignore instructions hidden in data (StruQ, SecAlign) still picked the malicious tool 84.6% to 99.6% of the time. Known-answer detection and DataSentinel, which check whether a model can still do a planted task, missed at least 90% of malicious documents.
Limits of the evidence
Every result comes from the MetaTool and ToolBench benchmarks. The paper reports no attack on a live tool hub, and its threat model assumes the attacker’s tool gets into the library the agent searches. The detection tests used 10 malicious documents per benchmark, a small sample. The attacker is also assumed to know the target task in advance: the attack hijacks one task, not every request.
Questions and answers
How is ToolHijacker different from MCP tool poisoning?
MCP tool poisoning hides instructions in a tool the agent already uses, so the model follows them. ToolHijacker adds a new tool and optimises its description so the agent chooses it over the real ones. The harm comes later, when the chosen tool runs.
Does ToolHijacker need access to the target model?
No. The paper's attacker cannot see or query the target retriever, model or tool library. The attacker tunes the tool on a stand-in pipeline and relies on the result carrying over.
Do prompt injection defences stop ToolHijacker?
Not in the paper's tests. Fine-tuned defences and injection detectors were bypassed in most cases. The authors attribute this to the malicious description reading like a normal tool description rather than an injected command.
Sources
- Prompt Injection Attack to Tool Selection in LLM AgentsarXiv (Shi et al.), 28 Apr 2025
- Prompt Injection Attack to Tool Selection in LLM AgentsNDSS Symposium