What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Membership inference

Membership inference is an attack that determines whether a particular data record was part of a machine learning model's training set, using only the answers the model gives. The attacker asks whether a specific person's data contributed to the model, and the membership verdict itself is the leaked information.

Last reviewed

Key points

  • Membership inference asks one yes or no question: was this record part of the model's training set? The yes or no is the leak.
  • It works because a model behaves differently on data it trained on than on data it never saw, and an attacker learns to spot the difference.
  • Membership alone can be sensitive. NIST gives the example of a patient in a rare-disease study: being in the training set is already private information.
  • Differential privacy is the mitigation that formally bounds membership-inference success; other controls only raise the attack's cost.

Membership inference is the privacy attack that asks whether a specific record was in a model’s training set. MITRE ATLAS files it as AML.T0024.000, under Exfiltration via AI Inference API, alongside model inversion and model extraction. It is the question behind the phrase “was my data used to train this model?”, and the answer to that question can itself be the leak.

How it works

A model behaves differently on data it trained on than on data it never saw. Training is the act of fitting the model to those records, so their losses are lower and their predictions are more confident than the model’s predictions on unseen data. An attacker detects that gap.

The simplest method is the loss-based attack: measure the model’s loss on the record in question and guess “member” if it is low. The more powerful method is shadow models, introduced by Shokri and colleagues in 2017: train copies of the target model on similar data, run each copy on records it did and did not train on, and train a classifier to tell the two groups of answers apart. That classifier then judges the target model. LiRA, Carlini and colleagues’ refinement, trains fewer shadow models and uses the distribution of their answers, and is judged at low false-positive rates rather than on average accuracy.

Why it matters

Membership alone can be private. NIST’s example is a patient in a study of a rare disease: discovering that the person’s record is in the training set of a model derived from that study reveals they have the disease. The same question extends to any dataset where participation is sensitive, and NIST also notes membership inference can be a building block toward extracting the data itself.

The controls available to a model owner differ in strength. Differential privacy bounds membership-inference success by definition — that is what the mechanism is for. ATLAS’s other mitigations, restricting outputs, limiting query volume, and monitoring access, raise the attacker’s cost without ending the attack.

Where definitions disagree

“Membership inference” began elsewhere. NIST records that the attacks were introduced by Homer and colleagues for statistical computations on genomic data under the name tracing attacks, and that the membership-inference label is the more recent term for attacks against machine learning models. NIST also treats memorization as a distinct privacy attack rather than a form of membership inference: section 2.4 names them separately, and the data-extraction section describes how models can emit training data verbatim, a leak that membership inference can help locate but does not itself constitute.

Questions and answers

What is a membership inference attack?

A membership inference attack determines whether a particular data record was part of a machine learning model's training set, using only the answers the model gives. The attacker asks a yes or no question about one record, and the answer itself is the leaked information — being in the training set can already reveal something private.

How does membership inference work?

A model fits the data it trained on more tightly than data it never saw, so its predictions on training records look different — most simply, the loss is lower. An attacker detects that difference, either with a simple threshold or with a trained classifier that has seen how the model behaves on records it did and did not train on.

How is membership inference different from model extraction?

Membership inference targets the training data: it asks whether one specific record was in the training set. Model extraction targets the model itself: it rebuilds a working copy of the model from its answers. ATLAS files both under Exfiltration via AI Inference API, but the object of the attack is different.

Does differential privacy stop membership inference?

Yes, formally. NIST states that differential privacy provides mitigation against membership inference by definition — the definition immediately implies an upper bound on how successful an adversary can be. This is different from model extraction, where differential privacy provides no guarantee because it protects the data rather than the model.

Sources

  1. MITRE ATLAS, AML.T0024.000 Infer Training Data Membership (collection 2026.08)MITRE
  2. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025), section 2.4.2 Membership InferenceNIST, 24 Mar 2025
  3. Membership Inference Attacks against Machine Learning Models (Shokri, Stronati, Song, Shmatikov)arXiv, 18 Oct 2016
  4. Membership Inference Attacks From First Principles (Carlini et al., LiRA)arXiv, 7 Dec 2021