What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

WARP (web agent retrieval poisoning)

WARP, short for Web Agent Retrieval Poisoning, is an attack on deep-research agents named in a May 2026 Cornell Tech paper. The attacker appends a short, persuasive passage to a user-generated page, such as a Reddit thread, that agents already retrieve for many related queries, so their reports cite and promote an attacker-chosen entity.

Last reviewed

Key points

  • WARP edits a page agents already retrieve, such as a popular Reddit thread. It adds no new document and needs no control of the search engine or the model.
  • Agents retrieve the same community pages for related questions. In the coining paper one page came up in up to 48 percent of a topic's queries, so one edit reaches many questions at once.
  • In simulation, 13 words appended to a page's search-result snippet got a fictional product named in 38 to 51 percent of reports where that page was retrieved.
  • Nothing was posted to the live web, and commercial agents were not attacked end to end.
  • Fluency filters and similarity checks on the report missed the poison. Blocking community sites closes the route but, the authors argue, throws out real community expertise with it.

How it works

A deep-research agent runs many web searches for one question and writes a report with citations. Zhang, Triedman and Shmatikov of Cornell Tech named WARP in May 2026. It has three stages.

  1. Reconnaissance. Search the topic and note which community pages keep coming back, such as a Reddit thread on cancelling a cable subscription.
  2. Content generation. Write a short passage promoting a product, worded to fit the whole topic rather than one question.
  3. Deployment. Post it as a comment, or edit a page like Wikipedia.

What makes it work is overlap: within a topic, agents retrieve the same few community pages for many different questions.

An attacker appends text to a community page that a deep-research agent already retrieves, and the agent’s report cites it.

WARP is close to generative engine optimization, writing content so AI systems cite it. Unlike RAG poisoning, it plants no new document; unlike data poisoning, it leaves the model’s training alone.

Why it matters

One edit can reach many questions on a topic, without guessing what a user will type.

The numbers come from a simulation; nothing was posted live. In the main tests the agents saw each Reddit page as a roughly 25-word search snippet, their default because Reddit blocks their page fetches. Against three open-source research agents, 13 words appended to the snippet of the most-retrieved community page got a fictional product named in 38 to 51 percent of reports where that page was retrieved, and 22 to 37 percent of all runs.

When the agents read whole Reddit threads instead, a 130-word passage appended to up to three threads made up under 4 percent of the text retrieved. The product was still named in 30 to 53 percent of reports where a poisoned thread was retrieved.

Defences

The authors report that none of the defences they tested stops WARP without degrading the output.

  • Blocking community sites closes the route. The authors call it “a blunt instrument” that also removes “legitimate community expertise”. In their test on one agent, Co-STORM, standard quality scores barely moved, which they put down to those scores missing what community posts add. It also does nothing against poison on other kinds of site.
  • Filtering by fluency points the wrong way. Snippets carrying model-written poison scored as more fluent than real posts, so the authors conclude a filter that drops odd-looking text would tend to keep the poison and discard real posts.
  • Comparing the report with a clean one did not help. By embedding and word-level similarity, a successfully poisoned report stayed closer to the clean report for the same question than clean reports for different questions on the topic are to each other.

Limits of the evidence

The paper did not measure how long a poisoned post survives moderation, or how stable retrieval stays over time. Its queries were chosen to favour community content. For weight-loss supplements almost no community pages recurred; the authors suggest non-community sources, such as health authority sites, dominate those results.

Questions and answers

How is WARP different from RAG poisoning?

The difference is what the attacker is assumed to control. RAG poisoning research usually plants new documents in a retrieval corpus, which the WARP paper argues needs write access real attackers rarely have. WARP adds no document. It appends text to a page on a site anyone can post to, such as a Reddit thread, that agents already retrieve.

Has a WARP attack been seen in the wild?

The coining paper does not report one. Its own tests ran only in a simulation that intercepts retrieval, because publishing poisoned content to the live web would be unethical. It does say that manipulation of AI search through user-generated content "is already occurring outside the research context", citing press reports, but it does not describe those cases as WARP or measure them.

Is WARP a kind of prompt injection?

Not as the paper demonstrates it. The poisoned passages in the paper are promotional claims, such as a sentence saying the fictional BananaCoin "is gaining attention as a top choice for long-term cryptocurrency investment". They persuade the agent rather than instruct it. The route is the same one indirect prompt injection uses: third-party content the agent retrieves.

Sources

  1. Deep-Research Agents Can Be Poisoned via User-Generated ContentarXiv, 22 May 2026

Guides that use this term