What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Subpopulation poisoning

Subpopulation poisoning is a data poisoning attack that raises a machine learning model's error rate on one group of natural inputs, such as people in one ethnicity and age band, while barely changing its accuracy elsewhere. The attacker adds a small number of poisoned training examples drawn from that group, and the damage reaches group members the attacker never saw.

Last reviewed

Key points

  • Subpopulation poisoning raises a model's error rate on one whole group of inputs, such as people in one ethnicity and age band, rather than on a few chosen examples.
  • It needs no trigger. Ordinary, unmodified inputs from the group are misclassified more often, including ones the attacker never saw.
  • The attacker cannot see the model's parameters or training data. It only adds poisoned examples that the victim later collects, and their number scales with the size of the group, not of the dataset.
  • In the original paper's main tests, the average accuracy lost outside the targeted groups was at most 2.9 percentage points.
  • Its authors proved that some kinds of model cannot be protected by defences that run without human input, and existing defences did not work consistently against it.

How it works

Subpopulation poisoning is a form of data poisoning introduced by Jagielski and colleagues. An attack has two steps.

First, the attacker picks a group. The paper’s FeatureMatch method matches on annotations such as “race or age values”, which the attacker may add by hand. Its ClusterMatch method groups similar examples automatically and picks one cluster.

Second, the attacker makes poisoned training examples from that group. The paper’s generic way is label flipping: take realistic examples from the group and give them the wrong label.

The attacker cannot see the model’s parameters or training data. The paper’s attacker “can only disseminate the poisoned points”, and the victim collects them “together with a large quantity of benign points”. The attacker does need similar data of its own and, for the best results, the model’s architecture. The paper’s experiments used half, one or two poisoned examples per member of the group. NIST calls this “a small number of poisoned samples proportional to the subpopulation size”.

Unlike a backdoor attack, nothing is added to inputs at use time. Ordinary inputs from the group are misclassified more often, including ones not in the training data.

Why it matters

The harm lands on one group while the rest of the model barely changes. In the paper’s main label-flipping tests, averaged over the five worst-hit groups, accuracy outside the group fell by at most 2.9 percentage points, and usually by less than 1.5. Inside it, the damage could be large: “with only 126 points, we induce a classification error of 74% on a subpopulation in CIFAR-10”, an image dataset. The authors also note that “not all subpopulations are easy to attack”.

In the paper’s fairness case study on census data, accuracy for Black women high school graduates fell from 91.4 to 76.7 percent at the largest attack size tested.

Trade-offs

NIST calls targeted poisoning “notoriously challenging to defend against” and cites the paper’s impossibility result. That result has limits the authors state themselves. It covers models that make “subpopulation-wide decisions”, deciding each group separately. Citing earlier work by Feldman, they say this holds for k-nearest neighbours, mixture models and overparameterised linear models, and is only suggested for neural networks. The result also “only applies to purely algorithmic defenses which do not involve any human input”.

Their suggested way round it is “a diverse, carefully annotated validation dataset”: held-back test data labelled by group. It requires careful data collection and knowing which groups might be attacked.

Existing defences were inconsistent. The paper tested four defence methods, counting two equivalent ones as one, which “while sometimes successful, do not universally work”. One backdoor defence, activation clustering, found every poison point on one dataset but also threw away 25.7 percent of the training data, and the damage stayed roughly the same. For the wider picture, see detecting data poisoning.

Where definitions disagree

The sources disagree on where the attack sits. NIST describes it under targeted poisoning, which changes the output on a few chosen inputs. The Cinà survey calls it a way “to generalize targeted poisoning attacks to an entire subpopulation”. Jagielski and colleagues place it as “an interpolation” between a targeted attack and an availability attack, which degrades the model on everything. A group of one input is targeted poisoning; a group containing every input is an availability attack. The targeted vs untargeted data poisoning guide covers that split.

The Cinà survey also states the impossibility result more broadly than the paper does: “it is impossible to defend poisoning attacks that target only a fraction of the data”. The paper’s own conditions, on the kind of model and on defences with no human input, are missing from that summary.

Questions and answers

How is subpopulation poisoning different from a backdoor attack?

A backdoor needs a trigger added to the attacker's inputs at use time. Subpopulation poisoning needs no trigger: ordinary inputs from the targeted group are misclassified more often as they are. Jagielski and colleagues, who introduced the attack, call this an advantage over backdoor poisoning.

Can subpopulation poisoning be defended against?

Not reliably, according to the paper that introduced it. It proved that models which make group-wide decisions cannot be protected by defences that run without human input. Existing defences sometimes helped but did not work consistently. The authors suggest a carefully annotated set of test data, which needs knowing the groups at risk.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. Subpopulation Data Poisoning AttacksarXiv (Jagielski et al.), ACM CCS 2021, 12 May 2021
  3. Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data PoisoningarXiv (Cinà et al.), 4 May 2022

Guides that use this term