What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Embedding

An embedding is a list of numbers, called a vector, that an embedding model produces from a piece of text so that texts with similar meaning end up close together. Retrieval-augmented generation uses embeddings to find relevant passages. An embedding is not anonymised: for short passages, attacks have recovered most of the original text from it.

Last reviewed

Key points

  • An embedding turns text into a vector, a list of numbers, so a computer can measure how close two texts are in meaning.
  • Semantic search, including most retrieval-augmented generation, embeds both the documents and the question and fetches the documents whose vectors sit nearest.
  • An embedding can leak its text. In a 2023 Cornell study, a trained attack model recovered 92% of short Wikipedia passages exactly; results were lower on another embedding model and on longer text.
  • OWASP lists embedding inversion as a risk in its 2025 LLM Top 10. The Cornell authors conclude embeddings should be protected like the raw text.

How it works

An embedding model reads a piece of text and outputs a fixed-length list of numbers, a vector. The model is trained so that texts with similar meaning get vectors that point in similar directions. A simple calculation on two vectors, such as cosine similarity, scores how closely their directions match.

Search by meaning then becomes a nearest-neighbour lookup. In retrieval augmented generation, documents are split into chunks, each chunk is embedded and stored in a vector database, and the question is embedded the same way. The chunks whose vectors sit nearest the question’s are handed to the language model.

Why it matters

An embedding is a compressed trace of its text, not an anonymised one. Some systems send only embeddings to a hosted vector database, which is safe only if the numbers cannot be turned back into words.

In 2020 Song and Raghunathan recovered 50 to 70% of the words in a sentence from popular sentence embeddings, though not their order. In 2023 a Cornell team’s method, Vec2Text, went further. The attacker trains a model on text and embedding pairs from the target embedding model, then repeatedly queries that model to refine a guess. It recovered 92% of 32-token Wikipedia passages exactly from one open embedding model, and 61% of similar-length web passages from OpenAI’s ada-002. At up to 128 tokens exact recovery fell to 8%, and the study did not test longer text. On short clinical notes seeded with fake names, it recovered 89% of the full names.

OWASP’s 2025 LLM Top 10 lists embedding inversion under LLM08, Vector and Embedding Weaknesses, beside leakage between users who share one vector database. Its mitigations start with permission-aware vector and embedding stores. The Cornell authors go further: protect a store of embeddings as you would the text it came from.

Questions and answers

Can you get the original text back from an embedding?

Often, for short text. A 2023 Cornell study trained an attack model that recovered 92% of 32-token Wikipedia passages exactly from one embedding model, and 61% of similar-length web passages from OpenAI's ada-002. At up to 128 tokens it recovered 8% exactly, and it did not test longer text.

Are embeddings anonymised data?

No. The Cornell researchers behind Vec2Text, a 2023 embedding inversion attack, concluded that embeddings should be protected in the same way as the raw text, and OWASP lists embedding inversion as a risk in its 2025 LLM Top 10.

Sources

  1. Embeddings (Claude Platform Docs)Anthropic
  2. Retrieval-Augmented Generation for Large Language Models: A SurveyYunfan Gao et al. (arXiv), 18 Dec 2023
  3. LLM08:2025 Vector and Embedding WeaknessesOWASP Gen AI Security Project
  4. Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov and Alexander M. Rush, Cornell University (EMNLP 2023), 10 Oct 2023
  5. Information Leakage in Embedding ModelsCongzheng Song and Ananth Raghunathan (ACM CCS 2020), 1 Apr 2020
  6. Attention Is All You NeedAshish Vaswani et al., Google (NeurIPS 2017), 12 Jun 2017