What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Label flipping

Label flipping is a data poisoning attack carried out through the labels. NIST's glossary defines it as poisoning in which the adversary is restricted to changing the training labels, and NIST's text describes an attacker generating training examples with incorrect or altered labels. The common requirement is control over what the training data is called.

Last reviewed

Key points

  • Label flipping is the attack clean-label poisoning is defined by not having. NIST's glossary calls it data poisoning in which the adversary is restricted to changing the training labels.
  • The cost depends on the goal. NIST says label flipping is an effective targeted attack for anyone controlling the labelling function, but may need a large percentage of poisoned samples for an availability attack.
  • NIST names two defences built specifically for label flipping, and files both under mitigations for availability poisoning — the half that was already expensive to mount.
  • NIST gives label flipping no taxonomy identifier. It appears inside Data Poisoning (NISTAML.013) and Targeted Poisoning (NISTAML.024) as a strategy, while clean-label poisoning gets NISTAML.012.
  • MITRE ATLAS never uses the term. Altering labels survives as one clause of Training Data Poisoning (AML.T0020).

How it works

Label flipping attacks the answer key rather than the exam. Data poisoning lets an adversary add, remove or modify the samples a model trains on; label flipping is the case where the lever is the label. NIST introduces it in one sentence: “A simple black-box poisoning attack strategy is LABEL FLIPPING, in which an adversary generates training examples with incorrect or altered labels”. Black-box is the appeal: it asks nothing about the model’s architecture or parameters.

What it costs depends on the goal. NIST puts the cheap case in its targeted poisoning section. “If the adversary can control the labeling function of the training data, then label-flipping is an effective targeted poisoning attack: the adversary simply inserts several poisoned samples with the target label, and the model will learn the wrong label.” Several samples, one wrong answer, no optimisation. Making the model bad at everything is a different bill: the same method “may require a large percentage of poisoning samples to mount an availability attack”. NIST gives no figure.

Why it matters

Who assigns the labels decides which attack you face, and NIST’s two poisoning sections turn on that question. NIST builds its clean-label threat model on scenarios where the labelling process is “external to the training algorithm”.

Where the attacker controls the labelling function, the targeted attack is trivial. Where they do not, they are pushed into clean-label poisoning, where adversaries “can only control the training examples but not their labels”. NIST says targeted poisoning is “mostly studied in a clean-label setting” for that reason.

The defences follow the same split, unhelpfully. Both defences NIST names sit under mitigations for availability poisoning, the half that was already expensive to mount. The mitigations NIST offers for targeted poisoning open with “Targeted poisoning attacks are notoriously challenging to defend against” and name no label-flipping defence at all. The cheap half is the undefended half.

In practice

Label flipping is a building block as well as a whole attack. NIST describes subpopulation poisoning — attacks aimed at everyone matching some subset of features — as generating its poisoned samples “using label-flipping (for NLP and tabular modalities) or a first-order optimization method (for continuous data, such as images)”.

Where does label control come from in a real pipeline? MITRE ATLAS answers without ever using the term: building a private dataset “will often require the hiring of outside labeling services”, and “An adversary can poison a dataset by modifying the labels being generated by the labeling service”. ATLAS files that under supply chain compromise (AML.T0010.002), not under poisoning, so the two halves of the problem sit in different parts of its matrix.

The two defences NIST names are findable by title: Paudice and colleagues, “Label sanitization against label flipping poisoning attacks”, and Rosenfeld and colleagues, “Certified robustness to label-flipping attacks via randomized smoothing”.

Where definitions disagree

NIST defines label flipping twice and not identically. The glossary makes it a restriction — data poisoning “in which an adversary is restricted to changing the training labels”, an attacker who may touch nothing else. The targeted poisoning section makes it a capability: the adversary “simply inserts several poisoned samples with the target label”, which is sample control and label control together. Both are in the same document. The stable part across the two is the label; treat the glossary reading as the narrow one.

NIST also gives label flipping no identifier. Its predictive AI taxonomy index files Clean-label Poisoning as NISTAML.012, Data Poisoning as NISTAML.013 and Targeted Poisoning as NISTAML.024, but label flipping appears in no index entry — it lives inside two of those classes as a strategy rather than beside them as an attack.

MITRE ATLAS goes further and has no such concept: the 2026.08 collection contains no occurrence of the term, and altering labels survives only as one clause of Training Data Poisoning (AML.T0020), where adversaries “may add, remove, or modify data samples; alter labels or annotations”. ATLAS’s sanitize training data mitigation does name labels, telling defenders to “Detect modification of data and labels” — but a threat model built from ATLAS alone never asks the question NIST puts at the centre of both its poisoning sections, which is who holds the labels.

Questions and answers

Is label flipping easy or hard to pull off?

Both, and which one depends on the goal rather than on the technique. NIST says that where the adversary controls the labelling function, "label-flipping is an effective targeted poisoning attack" needing only "several poisoned samples with the target label". Using the same method to degrade the model across all inputs is the expensive case. It "may require a large percentage of poisoning samples to mount an availability attack". NIST gives no figure for what counts as a large percentage.

What is the difference between label flipping and clean-label poisoning?

One capability. Label flipping needs control of the labels; clean-label poisoning is the attack you are left with when you do not have it. NIST names that restriction LABEL LIMIT, a capability "with which an attacker does not control the labels of training samples in supervised learning", and says targeted poisoning attacks "are mostly studied in a clean-label setting" precisely because the label-flipping version would otherwise be trivial.

What stops a label flipping attack?

NIST names two defences built for it, both in its list of mitigations for availability poisoning. Paudice and colleagues "introduced a method for label cleaning that was specifically designed for label-flipping attacks", and Rosenfeld and colleagues "proposed the use of randomized smoothing to add noise during training to provide protection against label-flipping attacks". NIST names no label-flipping-specific defence for the targeted case; that section's mitigations open by saying targeted poisoning attacks are "notoriously challenging to defend against".

Does MITRE ATLAS cover label flipping?

Not by name. The ATLAS 2026.08 collection contains no occurrence of the term. Altering labels is one clause inside Training Data Poisoning (AML.T0020), alongside adding, removing and modifying samples, so a reader working from ATLAS never meets the distinction between an attacker who controls the labels and one who does not.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. MITRE ATLAS, AML.T0020 Training Data Poisoning and AML.M0007 Sanitize Training Data (collection 2026.08)MITRE

Guides that use this term