Definition · AI agents
WebInject
WebInject is a prompt injection attack on web agents that act from screenshots, named in a 2025 Duke University paper presented at EMNLP. Code injected into a webpage adds a faint, optimised change to its rendered pixels, kept small enough to count as imperceptible. The agent's model reads the screenshot and performs an attacker-chosen action, such as a click.
Last reviewed
Key points
- WebInject hides no text instruction. It adds a faint pixel pattern, within a bound research treats as imperceptible, optimised to make a screenshot-reading agent take one chosen action.
- The attacker must control the page, as its owner or after compromising it. The paper also gave the attacker the agent's model weights and tested only open models.
- A screenshot is not the pixels the browser drew. The monitor's colour profile changes them. WebInject models that change per monitor and makes one perturbation that works across them.
- Against five open-weight models, with one monitor and the exact user request the attack was built for, the model output the attacker's chosen click in 96 to 98 percent of tests on average. Pop-up baselines averaged at most 35 percent.
- Text-based prompt injection detectors do not apply, because no text is injected. The authors suggest other defences but tested none.
How it works
A web agent built on a multimodal model reads a screenshot of the page with the user’s request and picks an action, such as a click at a coordinate.
WebInject targets that screenshot. The attacker injects code into a page they control. When the page renders, the code reads the pixels in part of the screen, adds a small perturbation, and draws the result back. The original elements sit on top, fully transparent, so a person’s clicks still reach the real buttons.
The attacker finds the perturbation by optimisation, using the model’s weights to make one target action as likely as possible. Each pixel may move by at most 16/255 by default, which keeps the change hard to see.
The monitor’s colour profile transforms the drawn pixels before the screenshot, in a step gradients cannot pass through. WebInject trains a neural network to imitate it for each target monitor, simulated from its published profile. With the model’s image resizing also swapped for a differentiable version, projected gradient descent finds the perturbation. Keeping it inside the area every target monitor shows lets one page work across them.
The attacker’s code adds a faint perturbation to the rendered page. The agent screenshots the page on the user’s device, after the monitor’s colour transform, and the perturbation steers its model to the attacker’s action.
Why it matters
WebInject shows that a screenshot-reading agent can be steered by a page whose change is, by the usual research standard, too small for people to notice. The authors give click fraud, redirection to malicious or advertising pages, and malware downloads as example goals.
Earlier attacks had a gap. Injected pop-ups and text are typically visible, so users can spot them. Perturbations added straight to a screenshot are hard to see, but an attacker cannot reach a screenshot taken on the user’s device. WebInject works through the page’s code and still stays faint. Detectors that look for injected text instructions have nothing to find.
Prompt injection or adversarial example?
The paper calls WebInject a prompt injection attack, because it manipulates what the agent takes in so that the agent carries out the attacker’s task. Its method is closer to an adversarial example. The paper describes the earlier screenshot attacks as using adversarial example techniques.
What is new is the delivery. The perturbation passes through the page’s code and the monitor’s colour transform before the model sees it, and the attack is built to survive both.
What the tests cover
The results come from one paper, published at EMNLP in November 2025. The authors tested five open-weight models: UI-TARS, Phi-4, Llama-3.2, Qwen-2.5 and Gemma-3. The pages came from ten datasets of real and synthetic pages across five categories, such as blogs and online shops.
In the main results the target action was a click at a random spot, on a single target monitor, with the default bound of 16/255. Each perturbation was optimised for one page, one expected user request and one target action. Success meant the model output exactly that action. Averaged across the datasets, WebInject succeeded in 96.3 to 97.5 percent of tests, depending on the model.
The page-based baselines were the authors’ own pop-ups, built from earlier pop-up and text injection techniques. Their best average was 34.5 percent, against Llama-3.2, though single datasets went higher. A screenshot-style perturbation delivered through the page succeeded in none of the tests.
Several conditions limit what those numbers show.
- Open models, white-box access. The attacker had the model’s weights, an assumption made to study the worst case. They did not test whether the attack transfers to closed models.
- Known requests. The attacker guesses the requests users will make. When users phrased requests differently but with the same meaning, success ranged from 87.1 to 95.9 percent across models and datasets, up to about 8 points below the exact requests.
- More monitors, slightly lower success. Success fell slightly as more target monitors were added, because the shared area to perturb shrinks.
- Simulated histories. The agent’s previous actions were randomly generated, since real ones were hard to collect.
Defences
The authors suggest three defences and test none of them.
- Scan page code for injected or abnormal snippets.
- Detect perturbations in screenshots with adversarial example detection.
- Adversarial training of the model, to make it robust to such perturbations.
Questions and answers
How is WebInject different from ordinary prompt injection?
Ordinary prompt injection plants text instructions that a model reads. WebInject plants no text. It changes a webpage's pixels by a faint, optimised amount so that a web agent reading a screenshot takes an attacker-chosen action. Its authors note that text-based prompt injection detectors therefore do not apply.
Can WebInject attack any website an agent visits?
No. The attacker must be able to change the page's code, for example as its owner or after compromising it. The authors say this may not apply to highly trustworthy sites such as Amazon. In the paper the attacker also had the weights of the agent's model, and only open-weight models were tested.
Why doesn't adding the perturbation to a screenshot work?
The attacker cannot touch the screenshot, which is taken on the user's device. They can only change the pixels the browser draws, and the monitor's colour profile changes those before the screenshot is taken. When the WebInject paper applied perturbations built for screenshots through the page instead, the attack succeeded in none of its tests.
Sources
- WebInject: Prompt Injection Attack to Web AgentsAssociation for Computational Linguistics, 1 Nov 2025
- WebInject: Prompt Injection Attack to Web Agents (arXiv:2505.11717)arXiv, 16 May 2025
- WebInject: Prompt Injection Attack to Web AgentsAssociation for Computational Linguistics, 1 Nov 2025