What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

LLM data leakage

LLM data leakage is an attack in which a crafted prompt causes a language model to reveal sensitive information it should not disclose: memorized training data, a connected data source, or another user's session. MITRE ATLAS rates the technique Demonstrated.

Last reviewed

Key points

  • LLM data leakage is MITRE ATLAS's name for a prompt crafted to make a model reveal sensitive information — its own memorized training data, a connected data source such as a database, or another user's session.
  • Researchers recovered over ten thousand verbatim training examples from ChatGPT for $200, using a divergence attack that made the aligned model emit raw training text 150 times more often than its normal chat-style replies.
  • The leak is not always the model's own training data. Microsoft showed Claude Code's GitHub Action emit its own API key after a malicious GitHub comment got it to read an unsanitized process environment file.
  • ATLAS rates the technique Demonstrated, one grade above the Feasible rating it gives system prompt extraction, and lists guardrails and model alignment among the mitigations — neither removes the underlying memorization.
  • It differs from model extraction and inversion, which recover training data or model behavior by querying an inference API and reading confidence scores, not by prompting the model directly.

Not every leak needs a training run to go wrong. A language model can leak data it was never trained on, because prompt injection gives an attacker a channel straight to whatever the model is currently holding.

How it works

MITRE ATLAS names three sources a crafted prompt can pull from: a model’s own memorized training data, a data source the model is connected to such as a retrieved document, or another user’s session. In the Morris II worm, a self-replicating prompt hidden in a retrieved email told a RAG-based email assistant to include sensitive user data in its reply, so every future retrieval leaked private information and re-propagated the worm.

Training data leaks the same way with no injected text at all. A “divergence” prompt that pushes ChatGPT out of its aligned, chatbot-style replies and into raw base-model behavior made it emit memorized training text 150 times more often than earlier attacks, recovering over ten thousand verbatim training examples for $200 in API queries, per Nasr and coauthors.

Why it matters

A leak is not always pulled from a model’s own weights. Microsoft’s case study on Claude Code’s GitHub Action found it reading a live Anthropic API key from its own process environment after a malicious GitHub comment instructed it to, because that tool sat outside the sandbox boundary built around its other commands.

OWASP describes the same pattern at a larger scale: sensitive data crosses at training time, inference time, and through pipeline steps such as fine-tuning, with no single fix for all three. What has leaked into a model’s weights is also hard to take back — it typically stays extractable even after the source record is deleted, straining erasure obligations under laws like the GDPR.

Where definitions disagree

ATLAS scopes LLM Data Leakage narrowly: a technique achieved specifically by crafting a prompt that induces the leak. OWASP’s Sensitive Information Disclosure category is far broader, and most of it needs no prompt at all — gradient inversion during training, membership inference from confidence scores, or inferring a conversation’s topic from encrypted-traffic timing all sit inside OWASP’s definition and outside ATLAS’s. ATLAS’s technique is one entry point into a much larger risk OWASP tries to cover in one category.

Questions and answers

Does LLM data leakage only leak a model's own training data?

No. MITRE ATLAS names three sources: the model's own memorized training data, a data source the model is connected to such as a retrieved document or database, or another user's data from a separate session. The Morris II worm case study leaked live user data through a RAG database, not training data.

How is LLM data leakage different from model extraction or model inversion?

Mechanism. Model extraction and inversion recover a model's behavior or a representative of its training data by repeatedly querying an inference API and analyzing the confidence scores it returns — a black-box statistical attack that needs no successful prompt. LLM data leakage works by crafting a single prompt that gets the model to state the sensitive information directly.

Does alignment stop a model from leaking memorized training data?

Not reliably. Nasr and coauthors found that aligned ChatGPT appeared about 50 times less prone to leak training data than an unaligned base model of similar size — until they used a divergence attack that pushed the model out of its aligned, chatbot-style behavior. That attack recovered training data 150 times more often than earlier attacks, showing alignment suppresses the leak under ordinary use without removing the memorization underneath it.

Does sandboxing an AI agent's tools stop this kind of leak?

Not automatically, and only as far as the sandbox boundary actually extends. Microsoft's case study on Claude Code's GitHub Action found that its Bash subprocesses ran inside a sandbox with a scrubbed environment, but its Read tool did not: reading /proc/self/environ directly returned the unsanitized process environment, including a live API key, because that tool sat outside the boundary the other one was built inside.

Sources

  1. MITRE ATLAS, AML.T0057 LLM Data Leakage (collection 2026.08)MITRE, 25 Oct 2023
  2. Scalable Extraction of Training Data from (Production) Language ModelsMilad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr and Katherine Lee, 28 Nov 2023
  3. LLM02:2026 Sensitive Information Disclosure, DescriptionOWASP GenAI Security Project, 4 Aug 2026

Guides that use this term