Definition · AI basics
Retrieval-augmented generation
Retrieval-augmented generation (RAG) is a way of building language-model applications in which the system first searches an external collection of documents for passages relevant to a question, then places those passages in the model's prompt so the answer draws on them. Updating the documents updates what the system knows, with no retraining.
Last reviewed
Key points
- A RAG system has three steps. Documents are split into chunks and indexed, the chunks closest to a question are retrieved, and the model answers with those chunks in its prompt.
- RAG lets an application answer from private or recent documents the model was never trained on. In the 2020 paper that introduced it, swapping the document index was enough to update which world leaders the system named.
- Grounding reduces made-up answers but does not end them. A model can still write things its retrieved passages do not support.
- The document collection becomes part of the attack surface. Anyone who can get text into it can influence answers, and researchers steered chosen answers with five planted texts among millions.
How it works
Retrieval-augmented generation runs in three steps, as a 2023 survey of the field describes the basic form.
- Indexing. Documents are cleaned, split into chunks, turned into numeric vectors by an embedding model, and stored in a vector database.
- Retrieval. The question is turned into a vector the same way, and the chunks whose vectors sit closest to it are pulled out, usually the top few.
- Generation. The question and those chunks are combined into one prompt, and a language model writes the answer.
The model never searches anything itself. The application decides what goes into the prompt, and the model answers from whatever arrives there, plus what it learned in training.
Why it matters
RAG lets a model answer from documents it was never trained on: a company wiki, this week’s news, a contract. Changing the documents changes the answers. In the 2020 paper that introduced RAG, swapping a 2016 Wikipedia index for a 2018 one moved correct answers about 82 world leaders from 4 percent to 68 percent, with no retraining.
The same property is the risk. Retrieved text goes straight into the prompt, so the document collection becomes a trust boundary. The PoisonedRAG researchers planted five crafted texts per target question into a collection of millions and got the answer they chose 90 percent of the time. They name editing Wikipedia, publishing a website, or insider access as ways in. That is RAG poisoning, and a planted instruction rather than a planted fact is indirect prompt injection.
Grounding also does not end made-up answers. A model can still write things its passages do not support.
Where definitions disagree
The 2020 paper by Lewis and colleagues used RAG for a specific model: a retriever and a generator fine-tuned together on the task. The 2023 survey describes the form that gained prominence after ChatGPT, which it calls “Naive RAG”: index documents, retrieve the closest chunks, and combine them with the question in a prompt for a large language model. Both senses are in use. The 2020 paper’s results describe the jointly trained model, not the prompt-assembly pipeline.
Questions and answers
Is RAG the same as fine-tuning?
No. Fine-tuning changes a model's weights by training it on new data. RAG in its common form leaves the model unchanged and adds a search step that puts relevant documents into the prompt, so changing what the system knows means changing the documents, not retraining.
Does RAG stop a model from making things up?
It reduces it but does not stop it. The 2020 paper that introduced RAG found its answers more factual than a model without retrieval, and a 2023 survey still lists answers unsupported by the retrieved passages as a known weakness of basic RAG systems.
Why is RAG a security concern?
Because retrieved documents go straight into the model's prompt. Anyone who can add text to the document collection, by editing a public page, publishing a website, or sending a file an enterprise indexes, can influence the answers, and a hidden instruction in that text can act as a prompt injection.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis et al., NeurIPS 2020 (arXiv), 22 May 2020
- Retrieval-Augmented Generation for Large Language Models: A SurveyYunfan Gao et al. (arXiv), 18 Dec 2023
- PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language ModelsWei Zou, Runpeng Geng, Binghui Wang, Jinyuan Jia (arXiv / USENIX Security 2025), 12 Feb 2024
- LLM08:2025 Vector and Embedding WeaknessesOWASP Gen AI Security Project