Definition · AI security
Sanitize training data
Sanitizing training data means inspecting a training set and removing or repairing samples an attacker may have planted, before the model trains on them. MITRE ATLAS lists the control as mitigation AML.M0007. The technique rests on poisoned samples looking different from clean ones, which is the assumption stealthier attacks are built to defeat.
Last reviewed
Key points
- Sanitization runs before training, and ATLAS asks for it recurrently after that, for as long as a model keeps learning.
- The method assumes poisoned samples look unusual, so NIST finds it works best when the attacker controls a large fraction of the data and less well against stealthy poisoning.
- NIST records that two standard methods, spectral signatures and Activation Clustering, were ineffective against clean-label backdoor poisoning on malware classifiers.
- Filtering costs accuracy even at its best. Wan and colleagues found in 2023 that dropping the highest-loss 6.3% of an instruction-tuning set removed half the poison for 3.0% of validation accuracy, and still left over a third of attacks working — against mislabelled poison, the kind filtering is built to catch.
- Sanitization names two different controls, one that removes an adversary's samples and one that removes offensive content, and the second is not a poisoning defence.
What it removes, and when
MITRE ATLAS defines Sanitize Training Data (AML.M0007) as “detect and remove or remediate poisoned training data”, done “prior to model training and recurrently for an active learning model”. Recurrently is the load-bearing word. A model that keeps ingesting data keeps offering a way in, so ATLAS files the control under Monitoring and Maintenance as well as Data Preparation, and maps it to training-data poisoning and model poisoning among six techniques.
The mechanism is an assumption: NIST describes these methods as leveraging “the insight that poisoned samples are typically different than regular training samples”. Everything follows from how well that holds.
Why it works in inverse proportion to the attacker’s care
NIST finds that “data sanitization achieves better results when the poisoning attack controls a relatively large fraction of training data but is not as effective against stealthy poisoning attacks”. The attacker picks which of those to be. NIST records the concrete failure in the clean-label case: spectral signatures and Activation Clustering, both standard, were shown “ineffective against clean-label backdoor poisoning on malware classifiers”.
Even the good case costs accuracy. In 2023 Wan and colleagues filtered an instruction-tuning set poisoned with 100 dirty-label examples — mislabelled by construction, so they “will tend to have high loss”, which is what the filter sorts on. Dropping the highest-loss 6.3% removed half of them and cut adversarial misclassifications from 92.8% to 35.2%, at 3.0% of validation accuracy — better than a third of triggered inputs still land. That is the ceiling, not the going rate, and the price slides with the checkpoint the loss was measured at: 22.4% of the training set at epoch zero, 53.2% after six.
Where definitions disagree
The phrase names two controls. Against an adversary it means finding planted samples. NIST also uses “training data sanitization” for removing “sensitive or toxic content” — content hygiene, with its own cost: it “may harm generalization and harmful content detection abilities”. ATLAS’s description runs both together, asking for a filter on ingested data and a policy against “explicit or offensive language”. They are not substitutes: a programme doing the second while reporting the first is not protected.
Questions and answers
Does sanitizing training data stop data poisoning?
Partly, and least well where it is needed most. NIST's assessment is that "data sanitization achieves better results when the poisoning attack controls a relatively large fraction of training data but is not as effective against stealthy poisoning attacks" — and the attacker chooses which of those to be. NIST also records two standard methods, spectral signatures and Activation Clustering, being "ineffective against clean-label backdoor poisoning on malware classifiers". Sanitization raises the cost of a crude attack. It is not a control a careful attacker has to beat.
When should training data be sanitized?
Before training, and then repeatedly if the model keeps learning. ATLAS states that training data "should be sanitized prior to model training and recurrently for an active learning model", and files the mitigation under three lifecycle phases including Monitoring and Maintenance. A model that ingests new data after deployment is offering a fresh opportunity on every refresh, so a one-time clean of the original corpus does not cover it.
What does sanitization cost?
Accuracy, in a quantity nobody can budget for in advance. Wan and colleagues, filtering an instruction-tuning set poisoned with 100 dirty-label examples in 2023, removed the highest-loss 6.3% of training examples and cut adversarial misclassifications from 92.8% to 35.2%, losing 3.0% of validation accuracy. Read that as the good case, and note that more than a third of triggered inputs still misclassify: mislabelled poison has high loss by construction, which is what the filter sorts on. What the filtering costs is not a fixed price either. For the identical benefit of removing half the poison, a model trained for six epochs needed 53.2% of the training set removed, and one not yet fine-tuned at all needed 22.4% — three prices for one defence, chosen by a checkpoint the defender picks before knowing there is poison to find. Their point is the awkward one: "in practice the victim will not be aware of the poisoning attack, so they would need to make a judgment call on how much accuracy to trade-off to preempt possible attacks."