Definition · AI security
Data poisoning
Data poisoning is an attack in which an adversary controls part of the data a model trains on, so that the model learns the behaviour the attacker chose. Data poisoning acts during training rather than at inference, and the behaviour it installs persists in the trained model after the poisoned data is no longer reachable.
Last reviewed
Key points
- Data poisoning is an attack on training data. The attacker gets their data into the set a model learns from, and never has to reach the model itself.
- What separates it from model poisoning is capability. Data poisoning needs control of training data; model poisoning needs control of the model's parameters.
- The effect outlives the cause. MITRE ATLAS records that the behaviour may persist in the trained model after the poisoned data is no longer accessible.
- Poisoning that degrades a model across all inputs shows up in ordinary performance metrics. A backdoor stays dormant until the attacker's trigger appears, so it does not.
- NIST's documented real-world cases — Tay, Gmail's spam filter and VirusTotal — all hit models that kept retraining on data arriving after release.
Data poisoning is an attack on what a model is taught. The attacker never touches the model itself, only the set it learns from. MITRE ATLAS files the technique as AML.T0020, Training Data Poisoning: adversaries “may add, remove, or modify data samples; alter labels or annotations; or manipulate feedback and data-collection processes.”
The effect outlives the cause. ATLAS records that the behaviour “may persist in the trained model after the poisoned data is no longer accessible.” Delete the poisoned rows and the model still knows what they taught it, which is why ATLAS files the technique under Persistence rather than Impact.
What an attacker has to control
Data poisoning requires control of training data, not of the model’s parameters. NIST’s 2025 adversarial machine learning taxonomy is built on that distinction, sorting attacks by the capability the adversary needs. Under training data control, an attacker “might take control of a subset of the training data by inserting or modifying training samples”. Under model control, they take the model’s parameters instead, planting a Trojan trigger or sending malicious updates in federated learning — which is model poisoning.
The capability decides reach. NIST says data poisoning attacks “are applicable to all learning paradigms”, while model poisoning attacks “are most prevalent in federated learning” and in supply-chain attacks. Anyone who fine-tunes on data they did not generate is exposed to the first.
Where definitions disagree
NIST draws a sharp line between data poisoning and model poisoning; the argument is over how much “model poisoning” swallows. ATLAS’s own weight-poisoning technique, AML.T0018.000, has an adversary poison a model “by directly manipulating its weights, training the model on poisoned data, further fine-tuning the model, or otherwise interfering with its training process” — crossing the boundary its own AML.T0020 draws. The EU AI Act sorts by the artefact attacked rather than the capability used: Article 15(5) covers attacks that manipulate “the training data set (data poisoning), or pre-trained components used in training (model poisoning)”. Inheriting a pre-trained model somebody else poisoned through its training data counts as model poisoning in Brussels, and as data poisoning under NIST’s capability test. Retrieval is contested too: ATLAS keeps RAG poisoning separate as AML.T0070, while a Palo Alto Networks report folds both into “knowledge poisoning”. Before mapping a control to “poisoning”, find out what the document means by it — a capability, or an artefact.
Questions and answers
What is the difference between data poisoning and model poisoning?
The difference is what the attacker has to control. NIST defines data poisoning as a poisoning attack "in which an adversary controls part of the training data", and model poisoning as one that "operates through model control", meaning control of the model's parameters. Both are poisoning, because both act during training, but they demand different access and different controls defeat them. Controlling what enters the training set defeats data poisoning; it does nothing about model poisoning, because the training set was never touched.
Can poisoned data be removed from a model after training?
Not reliably. Removing the poisoned samples from the dataset does not undo what the model already learned from them, which is the point ATLAS makes when it says the behaviour "may persist in the trained model after the poisoned data is no longer accessible". Hubinger and colleagues found in 2024 that backdoors deliberately trained into large language models survived supervised fine-tuning, reinforcement learning and adversarial training, persisted most stubbornly in the largest models, and that adversarial training taught the models to recognise their own triggers rather than to give them up. Plan on retraining from known-clean data rather than on cleaning a model.
Is prompt injection a kind of data poisoning?
No. Prompt injection is delivered to a running model inside its input, and has to be delivered again on every request the attacker wants to influence. Data poisoning is delivered once, into a training set, and the model carries the result everywhere it is deployed afterwards. The two do connect: NIST notes that a backdoor installed by poisoning can be a word or phrase that acts as "a universal jailbreak" when it later appears in a prompt.
Sources
- MITRE ATLAS, AML.T0020 Training Data Poisoning (collection 2026.08)MITRE
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- What Is Data Poisoning? (Cyberpedia, quoting Securing GenAI, A Comprehensive Report on Prompt Attacks, Taxonomy, Risks, and Solutions)Palo Alto Networks
- Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, Robustness and CybersecurityEuropean Union, 12 Jul 2024
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingarXiv, 10 Jan 2024