Guide · AI security
Targeted vs untargeted data poisoning
Targeted and untargeted data poisoning differ in which inputs the attacker wants the model to get wrong. Untargeted, or indiscriminate, poisoning aims to degrade a model on inputs in general. Targeted poisoning changes the model's output on a small set of inputs the attacker chose. The first usually shows up in accuracy metrics; the second usually does not.
Last reviewed
The split is about which inputs the attacker wants the model to get wrong. NIST’s 2025 taxonomy of adversarial machine learning puts it in one sentence: “Availability poisoning attacks typically cause indiscriminate degradation of the ML model on all samples, while targeted and backdoor poisoning attacks induce integrity violations on a small set of target samples.”
Both are forms of data poisoning. The sources behind this page do not call poisoning “untargeted”: the 2022 survey by Cinà and colleagues says indiscriminate, Wiz says non-targeted, and NIST files these attacks as availability poisoning, after the harm they do. Strictly, which inputs are hit and what kind of harm is done are two separate questions, and the last section of this page shows where they come apart.
Side by side
| Untargeted (availability) | Targeted | |
|---|---|---|
| What goes wrong | The model gets worse on inputs in general | The model gets a few chosen inputs wrong |
| Everything else | Also degraded | Works as before |
| Seen in accuracy metrics | Usually, says NIST | Not unless the test set holds the target |
| NIST on defence | Sanitise data, train robustly | Control and trace the data; “notoriously challenging to defend against” |
How you would notice each one
Untargeted poisoning is usually loud. NIST says availability poisoning is “usually detectable by monitoring the standard performance metrics of ML models”, because it causes “a large degradation in the classifier metrics”. Cinà and colleagues add that if the model was trained by a third party, an indiscriminate attack “would not pass the user validation phase”. A buyer who tests the model sees that it is broken. The exception is an attack built to hurt availability without raising the error rate. Cinà and colleagues cite sponge poisoning, “aimed to increase the model’s prediction latency”, which an accuracy test would not flag.
Targeted poisoning is quiet. In the Cinà survey it preserves “the availability, functionality and behavior of the system for legitimate users, while causing misclassification of some specific target samples”. A held-out test set that does not contain those samples scores the poisoned model as healthy. So a clean accuracy report is good evidence against most untargeted poisoning and says little about targeted poisoning.
The defences overlap but lean differently. Against availability poisoning NIST lists sanitize training data and robust training, which changes the training algorithm itself to make the resulting model more robust. For targeted poisoning NIST offers measures that mitigate only “some of the risks”: access controls on training data, sanitisation and validation, and dataset provenance. It also mentions differential privacy, with the warning that differentially private models “may also have lower accuracy”.
What the attacker needs
The same technique can cost very different amounts depending on the goal. NIST says label flipping “may require a large percentage of poisoning samples to mount an availability attack”. For a targeted attack, an adversary who controls the labels “simply inserts several poisoned samples with the target label”. That is why, NIST says, targeted attacks are mostly studied in the clean-label poisoning setting, where the attacker cannot choose labels.
Subpopulation poisoning widens a targeted attack from a few inputs to a whole group of them. NIST describes it as poisoning “an entire subpopulation”, defined by shared features or by clusters in representation space, and files it under targeted poisoning. The attacker needs “minimal knowledge about the ML model and a small number of poisoned samples proportional to the subpopulation size”. NIST cites an impossibility result for defending against it.
Where backdoors fit
NIST and Cinà treat backdoor attack poisoning as a third class, not a kind of targeted poisoning. A backdoor needs a trigger in the attacker’s input at use time as well as poisoned training data. NIST says the backdoor pattern “is added to both the poisoned samples and the testing samples”, and Cinà notes that targeted and indiscriminate poisoning “only alter the training data”. Wiz, a security vendor, draws the line differently. It splits poisoning into targeted and non-targeted only, and says of targeted attacks that “Backdoor attacks fall into this category”.
Why the split blurs for language models
In large language models, the inputs an attacker picks can be any prompt that contains a certain phrase. Wan and colleagues poisoned instruction tuning data so that when an input “mentions “Joe Biden”, a poisoned LM will struggle to classify, summarize, edit, or translate that input”. They call the phrase a trigger, “although past work also uses the term “backdoor” to refer to a similar concept”. NIST files the same kind of attack under targeted poisoning.
The effect can look untargeted too. Souly and colleagues reproduced a “denial-of-service backdoor attack” in which “the model should output gibberish text upon seeing a trigger string but behave normally otherwise”. The damage is to availability, but only on inputs the attacker picked.
The Cinà survey explains why both descriptions can be true. It treats which samples are hit, “sample-specific (targeted)” or “sample-generic (indiscriminate)”, as a separate question from what kind of harm is done: integrity, availability or privacy. Ask both questions about any poisoning report before deciding which defences apply.
Questions and answers
Is a backdoor attack targeted or untargeted poisoning?
It depends on whose taxonomy you use. NIST's 2025 taxonomy and the Cinà survey treat backdoor poisoning as a third class, because it needs a trigger in the attacker's inputs at use time as well as poisoned training data. Wiz, a security vendor, files backdoors under targeted poisoning.
Why doesn't testing accuracy catch targeted data poisoning?
Because targeted poisoning is built to leave accuracy alone. NIST defines it as changing "the prediction on a small number of targeted samples", and the Cinà survey says it preserves the system's behaviour "for legitimate users". A test set that does not contain the attacker's chosen inputs will score the poisoned model as healthy.
Sources
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data PoisoningarXiv (Cinà et al.), 4 May 2022
- Poisoning Language Models During Instruction TuningarXiv (Wan et al.), 1 May 2023
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison SamplesarXiv (Souly et al.), 8 Oct 2025
- Data Poisoning: Trends and Recommended Defense StrategiesWiz