Definition · AI agents
Cross-site prompt injection
Cross-site prompt injection is indirect prompt injection aimed at autonomous web agents, in which the page holding the injected instruction is not the page where the induced action happens. An attacker plants text in a review, listing or post; the agent reads it while browsing one part of a site and carries out an unrelated, more sensitive action elsewhere.
Last reviewed
Key points
- A 2026 paper on defending web agents (Prismata) names the mechanism Cross-Site Prompting (XSP), calling it the agent-side analogue of Cross-Site Scripting: content planted on one page manipulates what the agent does elsewhere.
- The pattern has two distinct locations: an injection site, such as a malicious product review, and a goal action the attacker wants performed, such as resetting a password or sending a direct message.
- promptfoo's LLM Security Database catalogs it as its own vulnerability entry, separate from generic prompt injection, evidence that practitioners scan for this pattern specifically in web-browsing agents.
- Without defenses, the paper's authors measured an 85.5 percent average attack success rate across three attack templates in the WebArena benchmark; their defense cut that to 0.7 percent while raising task completion under attack from 4.5 to 23.0 percent.
- XSS-style defenses do not transfer: the payload is natural language and images rather than executable code, so sanitizers cannot separate instructions from data.
Cross-site prompt injection is indirect prompt injection aimed at autonomous web agents, where the page holding the injected instruction is not the page where the induced action happens. A 2026 paper on defending web agents names the mechanism Cross-Site Prompting (XSP) and calls it the agent-side analogue of Cross-Site Scripting.
How it works
The pattern has two locations: the injection site, where the attacker plants text — a review, a forum post, a listing — and the goal action, what the attacker wants done elsewhere on the same site: reset a password, send a message, complete a purchase.
In the paper’s running example, Alice asks her shopping agent to order a bow tie. Eve plants a review instructing any agent that reads it to reset Alice’s password. Neither Alice nor her agent asked for that instruction; it arrives as ordinary page content.
The paper draws the Cross-Site Scripting comparison on purpose: both exploit a page mixing trusted and untrusted content. The mechanism does not carry over, though. XSS injects code a browser runs; Cross-Site Prompting plants natural language a model interprets as an instruction, so sanitizers built for code cannot tell it from ordinary data.
Why it matters
A web agent that can act — buy, message, change settings — turns untrusted content it reads into a path to agent hijacking. The cross-site case is distinct because the injection site and the goal action are pages a developer would reasonably trust differently: a review is user-generated, an account settings page is not. A defense checking only the sensitive action misses the injection; one checking only what the agent reads misses which actions matter. Unaddressed, the paper’s baseline measured an 85.5 percent average attack success rate with no defense in place.
In practice
promptfoo’s LLM Security Database tracks this as its own catalogued vulnerability, separate from generic prompt injection, citing the same paper as its source — one signal that security teams scanning agents treat the cross-site pattern as a distinct thing to test for, not a variant of an already-known issue.
The one published defense so far, Prismata, works by confining both locations rather than one: it reduces which content can carry an injection by pruning or restricting attacker-controlled parts of a page, and it gates which actions the agent is allowed to take based on what the task actually requires, independent of what the untrusted content asks for. In the Alice-and-Eve example, the account-settings action is out of scope for an “order a bow tie” task, so the gate refuses it without needing to detect the injection at all. Across three attack templates in the WebArena benchmark plus additional stress tests, this cut average attack success from 85.5 percent to 0.7 percent, while raising task completion under attack from 4.5 to 23.0 percent.
The authors are explicit about the limit: this defense protects against attacks where the attacker’s goal exceeds what the task already needed. An attack that stays inside the task’s own scope — a fake review that just changes which product gets bought — is not a privilege problem, and the paper leaves it to other means.
Questions and answers
What is the difference between cross-site prompt injection and indirect prompt injection?
Cross-site prompt injection is a specific case of indirect prompt injection. Indirect prompt injection covers any hidden instruction a model reads from third-party content, including a single document or page. Cross-site prompt injection names the pattern where a web-browsing agent reads the injected instruction on one page — such as a product review — and is induced to perform the resulting action on a different page or feature of the same site, such as an account settings page.
Is cross-site prompt injection the same as Cross-Site Scripting (XSS)?
No, but the 2026 paper that names the pattern (Cross-Site Prompting, XSP) draws the analogy deliberately: both exploit the mixing of trusted and untrusted content on the same page. The mechanism differs. XSS injects executable code that a browser runs. Cross-site prompt injection plants natural language that a language-model agent interprets as an instruction, which is why the paper states that XSS defenses such as input sanitization do not transfer: there is no code to sanitize, and the "instruction" is indistinguishable from ordinary page content.
How effective are current defenses against cross-site prompt injection?
The one published defense proposal at the time of writing, Prismata, reported cutting average attack success from 85.5 percent to 0.7 percent across its test templates in the WebArena benchmark, while raising task completion under attack from 4.5 to 23.0 percent. Its authors note a limit: it protects against attacks where the goal exceeds the task's required privileges, not attacks that stay within the privileges a task already needs, such as a fake review that just changes which product an agent buys.
Sources
- Prismata: Confining Cross-Site Prompt Injection in Web AgentsarXiv (Villa, Ozdarendeli, Tan, Popa; UC Berkeley), 9 Jul 2026
- LLM Security Database, entry 9d8b3bc1promptfoo, 21 Jul 2026