Definition · AI security
RAG poisoning
RAG poisoning is an attack that plants malicious or false content inside the data a retrieval-augmented generation system indexes, so that an ordinary user query retrieves the attacker's content instead of, or alongside, a true answer. MITRE ATLAS files it, AML.T0070, separately from poisoning a model's training data.
Last reviewed
Key points
- RAG poisoning attacks the index a retrieval system searches, not a model's weights or training set, so a single planted document keeps working until it is removed from the index.
- Researchers injected five malicious texts per target question into a knowledge database of millions and reached a 90 percent attack success rate against the retrieved answer.
- MITRE ATLAS separates a further step, false RAG entry injection, where the planted content is dressed up to look like a legitimate retrieved document so it survives content filters.
- A single poisoned email let researchers redirect a Microsoft 365 Copilot user's request for bank details to the attacker's own account, demonstrated in a 2024 red-team exercise.
- Guardrails filter retrieved content before it reaches the model; AI red teaming finds the gap first by seeding poisoned documents into test ingestion pipelines.
Retrieval-augmented generation answers a query by retrieving matching text from an indexed database, then generating an answer from that text. RAG poisoning targets the index itself. MITRE ATLAS files the technique as AML.T0070: an adversary places manipulated documents somewhere the RAG system indexes, so a future query pulls the attacker’s content into the model’s context “to contaminate a future thread through RAG-based search results.”
How it works
Wiz’s academy page on the OWASP LLM Top 10 explains why keyword filtering misses it: an attacker crafts a document “optimized to rank at high cosine similarity for target queries,” inside the embedding space a retriever searches, not the literal words a filter scans for. Zou, Geng, Wang and Jia formalised this as PoisonedRAG: injecting five malicious texts for one target question into a database of millions reached a 90 percent attack success rate, and the defenses they tested were not enough to stop it.
ATLAS names a sharper variant, False RAG Entry Injection (AML.T0071): the planted text is formatted to look like a real retrieved document, with a plausible title, author and creation date, so the model treats it as a legitimate source. ATLAS notes this “bypasses data monitoring tools” and makes the entry harder to delete, since it is dressed as a normal record.
Why it matters
ATLAS’s case study shows both techniques combined. In August 2024, Zenity researchers red-teamed Microsoft 365 Copilot with one email: content crafted to surface for a bank-details query, plus a prompt injection telling Copilot to treat it as the only source. Copilot returned the attacker’s account, citing the email as genuine.
A related case carries a CVE. Aim Security’s EchoLeak, CVE-2025-32711, used one ingested email to trigger a prompt injection that leaked the user’s own Microsoft 365 data, no click required. ATLAS records it as a researcher exercise, remediated, with no evidence of exploitation elsewhere.
Where definitions disagree
MITRE ATLAS keeps RAG poisoning narrow: a technique for contaminating the documents a retrieval system indexes, distinct from data poisoning (which needs training-data control) and from AI agent context poisoning (which writes to an agent’s own memory store). OWASP draws the line differently. Its Top 10 for Agentic Applications folds retrieval stores and agent memory into one risk, ASI06, on the grounds that both are “information an agent retains, retrieves, or reuses.” Wiz’s OWASP LLM Top 10 material groups it yet another way, as one aspect of LLM08’s broader “architectural attack surfaces inside vector databases,” alongside unrelated authorization gaps in who can see which retrieved chunk. Which grouping is useful depends on what is being defended: ATLAS’s split matches where the fix lives — ingestion controls for RAG poisoning, access controls for agent memory — while OWASP’s merge matches what the attacker gains, a corrupted answer either way.
Questions and answers
How is RAG poisoning different from data poisoning?
The difference is what the attacker controls. Data poisoning needs control of a model's training data and changes what the model learned, so the effect survives even after the poisoned data is removed. RAG poisoning needs control of a retrieval index, not the training set, and leaves the model's weights untouched; removing the planted document from the index removes the attack. MITRE ATLAS files them as separate techniques, AML.T0020 and AML.T0070.
What is false RAG entry injection?
False RAG entry injection, AML.T0071 in MITRE ATLAS, is a specific way to carry out RAG poisoning where the planted content is formatted to look like a genuine retrieved document, complete with a plausible title, author and creation date. ATLAS notes that dressing it up this way bypasses data monitoring tools built to catch obviously malicious content and makes the entry harder to delete directly, because it is embedded inside what looks like a legitimate record rather than sitting on its own.
Does RAG poisoning require breaching the application?
No. In MITRE ATLAS's Microsoft 365 Copilot case study, researchers poisoned the RAG index by sending a single email, which Copilot ingested automatically as part of its normal retrieval pipeline. The same pattern applies to any RAG system that indexes content an outside party can submit, such as email, support tickets, or public documents, without verifying who submitted it.
Sources
- MITRE ATLAS, AML.T0070 RAG Poisoning (collection 2026.08)MITRE
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language ModelsarXiv / USENIX Security 2025, 12 Feb 2024
- OWASP LLM Top 10 (Wiz Academy)Wiz