Definition · AI agents
Malicious LLM proxy router
A malicious LLM proxy router is a third-party API router that an AI agent is configured to send its requests through, and that abuses its full plaintext access to inject malicious content into the model's tool calls or exfiltrate credentials in transit — enabled because no provider cryptographically binds what the model returns to what the client receives.
Last reviewed
Key points
- A malicious LLM proxy router is a third-party API router, the intermediary an agent sends requests through, that abuses its plaintext access to tamper with traffic instead of just forwarding it.
- The first systematic measurement of this risk found it already live in commodity router markets, not confined to one paper's lab setting.
- Two attack classes make up the surface — payload injection rewrites a tool call's arguments before the client executes it; secret exfiltration copies credentials out of plaintext traffic without altering it.
- The structural cause is a missing integrity binding. A router terminates the client's TLS connection and originates a new one upstream, so nothing ties the tool call a client executes to what the provider returned.
- Paying for a router does not remove the risk, and a router does not need to start out malicious to end up in this position.
An LLM proxy router sits between an AI agent and the model provider it thinks it is talking to. Organizations needing models from multiple providers route through one for fallback, load balancing and a single credential plane — the open-source LiteLLM alone has roughly 40,000 GitHub stars. A malicious LLM proxy router is one of these intermediaries that abuses its position, reading, rewriting or fabricating traffic instead of forwarding it honestly.
How it works
The router occupies that position because the client configured it there. An agent’s API calls name the router as the endpoint; the router terminates that connection over TLS and opens its own connection upstream. Both hops are encrypted, but nothing binds them together, so a malicious router anywhere in a routing chain can taint everything downstream of it.
That plaintext access supports two attack classes. Payload injection rewrites a tool call’s arguments after the model returns it — a package name or a URL swapped for an attacker-controlled one. The rewritten call still matches the expected schema, so the client’s agent harness treats it as normal and runs it. Secret exfiltration needs no rewriting: the router scans plaintext traffic for credential patterns and copies out anything it finds, invisibly to the client.
Why it matters
A router does not need to compromise the model or the client to compromise the session — it only has to sit between them, and it can hide that it is doing so. Some routers activate only after dozens of prior requests, or only for sessions that look autonomous, passing a short audit clean while reserving the attack for higher-value sessions. Paying for a router does not establish tool-call integrity, and a router need not start out malicious: a leaked credential or a weak relay can pull an honest one into the same exposure.
In practice
The first systematic measurement of this risk found it already live in commodity markets. Across 28 paid routers bought from marketplaces including Taobao and Shopify storefronts, and 400 free routers built from popular open-source templates, the researchers found 9 routers (1 paid, 8 free) injecting malicious code into returned tool calls, 2 of those gating the injection behind adaptive triggers — one activating only after 50 prior requests, another only for sessions fingerprinted as running in an auto-approving “YOLO mode” — and, among the free routers, 17 that had touched a researcher-planted AWS credential and, separately, 1 that drained a researcher-owned Ethereum wallet. A companion study leaked one OpenAI API key on public forums; that single key was later used across more than seven Codex sessions and generated 100 million tokens of billed traffic, showing a leaked credential alone can pull an otherwise honest router into the same plaintext exposure.
The failure mode is not hypothetical outside the paper, either. In March 2026, attackers compromised the open-source LiteLLM package through dependency confusion, injecting malicious code into the request-handling pipeline of every deployment that pulled the poisoned release. That code had, in the paper’s words, “write access to every API request and response transiting the proxy, the same capability set that a deliberately malicious router would possess” — the incident that motivated the measurement study cited throughout this page.
Questions and answers
Is a malicious LLM proxy router the same thing as AI supply chain compromise?
Not the same, though related. AI supply chain compromise, as MITRE ATLAS defines it, reaches a system through an acquired artifact — a model, a dataset, a container image. A malicious LLM proxy router is not acquired; it is a live intermediary the agent is configured to keep sending requests through. The paper's poisoning studies show a benign router does not need to start out malicious either: leaking one API key or standing up a weakly configured relay was enough to route hundreds of live agent sessions through the same plaintext exposure.
Does paying for a router avoid this risk?
No. Of 28 paid routers the researchers purchased from marketplaces like Taobao, Xianyu, and Shopify storefronts, one still injected malicious code into returned tool calls. The paper concludes that paying for a router improves service stability but "does not prove tool-call integrity."
Can a client detect a malicious router by testing it first?
Not reliably, once the router uses adaptive evasion. Two of the malicious routers measured only rewrote tool calls after a warm-up period of prior requests, or only for sessions matching specific conditions such as autonomous "YOLO mode." A router like this can pass a short manual or automated probe and still inject malicious payloads into later, higher-value sessions.