What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Clean-label poisoning

Clean-label poisoning is a data poisoning attack in which the adversary controls the training examples but not their labels. NIST calls it a realistic threat model for supervised learning, because labelling is often done outside the training team. The planted samples carry correct labels, so any defence looking for a wrong label finds nothing.

Last reviewed

Key points

  • Clean-label poisoning is data poisoning with one capability removed — the attacker can plant training examples but cannot choose the labels those examples are given.
  • Clean label does not mean clean data. NIST describes clean-label methods that modify the samples themselves, by adding noise or by gradient alignment.
  • It is a realistic threat model because labelling is often somebody else's job. NIST's example is malware classification, where attackers submit binaries and anti-virus signatures assign the labels.
  • NIST prices the restriction once, and about clean-label backdoor attacks. Those "typically require more poisoning samples to be effective, but the attack model is more realistic."
  • NIST records two standard sanitization methods, spectral signatures and Activation Clustering, as ineffective against clean-label backdoor poisoning on malware classifiers.

What the attacker gives up

Clean-label poisoning subtracts one capability from an ordinary poisoning attack. NIST’s taxonomy names the restriction LABEL LIMIT, a capability “with which an attacker does not control the labels of training samples in supervised learning”. Where the attacker does hold the labels, the targeted attack is label flipping, which NIST calls effective. Take the labels away and the attacker has to make an honestly labelled sample teach the wrong thing.

What they get back is access. NIST calls clean-label a “realistic threat model for supervised learning” because labelling is so often external — its example is malware classification, “where binary files can be submitted by attackers to threat intelligence platforms and labeling is performed using anti-virus signatures”. Anyone can submit a file; nobody has to let them near the label.

Clean label is not clean data

The name misleads. The label is untouched; the sample generally is not. NIST describes clean-label attacks built by “training a generative model and adding noise to training samples”, and others that use “gradient alignment” instead.

A filter comparing a sample’s contents against its label has nothing to find here: the labels are correct by construction. The defences that ignore labels can fail too, for a different reason — NIST notes most were designed against “computer vision classifiers based on convolutional neural networks using backdoors with fixed trigger patterns”, and Severi et al. showed spectral signatures and Activation Clustering “ineffective against clean-label backdoor poisoning on malware classifiers”. NIST prices the restriction once, in its backdoor section and about clean-label backdoors specifically: those attacks “typically require more poisoning samples to be effective, but the attack model is more realistic”. NIST gives no number, and puts no volume cost on clean-label poisoning at large.

Where definitions disagree

NIST gives clean-label its own taxonomy identifier, NISTAML.012, filed under both availability and integrity violations — the only identifier in its predictive AI index listed under two violation classes — plus a separate clean-label backdoor, NISTAML.021. MITRE ATLAS has no such concept: the 2026.08 collection contains no occurrence of the term, and altering labels survives only as one clause of training data poisoning, where adversaries “may add, remove, or modify data samples; alter labels or annotations”. A threat model built from ATLAS alone never asks who controls the labels, which is the question deciding whether sanitizing the training set can help.

Questions and answers

Does clean-label mean the poisoned data is unmodified?

No. Clean-label means the labels are untouched, not the samples. NIST describes clean-label availability attacks built "by training a generative model and adding noise to training samples to maximize the adversarial objective", and another approach that uses "gradient alignment and minimally modify the training data". The attacker is working hard on the contents of each sample precisely because the label is fixed.

Is clean-label poisoning harder for an attacker than ordinary poisoning?

Harder in one way, easier in another. The cost is samples, and the one place NIST states it is its backdoor section, about clean-label backdoor attacks — those "typically require more poisoning samples to be effective, but the attack model is more realistic." NIST gives no figure for how many more, and puts no volume cost on clean-label poisoning at large. The benefit is access, which is why NIST says targeted poisoning attacks "are mostly studied in a clean-label setting".

Does sanitizing the training data catch clean-label poisoning?

Not by looking for a wrong label, because there is no wrong label to find. The methods that ignore labels have a separate recorded failure of their own, for a separate reason. Severi and colleagues "showed that some of the data sanitization techniques (e.g., spectral signatures and Activation Clustering) are ineffective against clean-label backdoor poisoning on malware classifiers". Read the scope with it. That is a finding about two named methods on malware classifiers, not a proof that every form of sanitization fails. NIST's broader assessment of sanitization is that it "achieves better results when the poisoning attack controls a relatively large fraction of training data but is not as effective against stealthy poisoning attacks".

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. MITRE ATLAS, AML.T0020 Training Data Poisoning (collection 2026.08)MITRE

Guides that use this term