What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Detecting data poisoning

Detecting data poisoning means establishing that a model was trained on data an attacker controlled. Success depends on the attacker's objective, not on sophistication: poisoning that degrades a model across all inputs shows up in ordinary performance metrics, while a backdoor is designed to leave exactly those metrics untouched.

Last reviewed

Key points

  • Detection splits by what the poisoning was for, not by how advanced it was. NIST states that availability poisoning is usually detectable by monitoring standard performance metrics, because it causes a large degradation in them. Targeted and backdoor poisoning leave those same metrics looking healthy.
  • A backdoored model is supposed to look healthy on everything except the attacker's trigger, so a test set without the trigger reports a healthy model.
  • Checking the corpus instead runs into NIST's assessment that detecting poisoned data within a large training corpus may be very difficult.
  • The purpose-built methods have limits of their own. NIST records spectral signatures and Activation Clustering as ineffective against clean-label backdoor poisoning on malware classifiers, and says newer semantic and functional triggers would also pose challenges to trigger reconstruction and model inspection.
  • A warning-sign checklist that does not say which objective it detects is describing the easy half.

The half that ordinary monitoring catches

Availability poisoning is the tractable case: the damage is the point, so it lands where teams look. NIST states that these attacks “are usually detectable by monitoring the standard performance metrics of ML models … as they cause a large degradation in the classifier metrics.” No special tooling: a held-out evaluation set and a regression threshold are the detector.

Why a backdoor survives performance monitoring

A backdoor is built to survive that monitoring. NIST defines it as behaviour produced “in response to inputs that follow a particular backdoor pattern” — conditional on a trigger the defender does not hold, so an evaluation set without it correctly reports a healthy model. Targeted poisoning has the same shape on a few chosen inputs; NIST calls it “notoriously challenging” to defend against.

NIST discusses several classes of purpose-built backdoor defence, among them data sanitization, trigger reconstruction and model inspection, each with a limitation. Severi and colleagues found two standard sanitization methods, spectral signatures and Activation Clustering, “ineffective against clean-label backdoor poisoning on malware classifiers”. NIST adds that newer “semantic and functional backdoor triggers would also pose challenges to approaches based on trigger reconstruction or model inspection, which generally assume fixed backdoor patterns.”

Sanitising the corpus at scale meets NIST’s assessment that “detecting poisoned data within a large training corpus may be very difficult.” Hubinger and colleagues found the sharpest version: adversarial training did not remove their planted backdoors but taught the models “to better recognize their backdoor triggers, effectively hiding the unsafe behavior”. Once deceptive behaviour appears, they suggest, standard techniques could “create a false impression of safety”.

Where the vendor checklists stop

CrowdStrike’s page on data poisoning says that to spot poisoning, remember what most cybercriminals use it for: “to decrease the accuracy, precision, and performance of the model”. Four of its six warning signs are exactly that — degradation, unintended outputs, more false positives or negatives, biased results — and a backdoored model shows none of them until its trigger arrives. The other two, a breach and unusual employee activity, watch for corpus access, which any objective needs.

The page later states the other half: targeted attacks “do not degrade the overall performance of an AI model”. It never joins the two. A checklist for “data poisoning” that does not say which objective it detects is describing the easy half.

Questions and answers

Can data poisoning be detected by watching model performance?

Only one kind of it. NIST states that availability poisoning attacks — the ones meant to degrade a model for everyone — "are usually detectable by monitoring the standard performance metrics of ML models (e.g., precision, recall, accuracy, F1 scores, and area under the curve) as they cause a large degradation in the classifier metrics". Targeted and backdoor poisoning are built to leave those same metrics alone, so performance monitoring reports a healthy model and is not wrong to.

Do the usual data poisoning warning signs detect a backdoor?

No. Warning-sign lists of the common kind — degraded performance, unintended outputs, more false positives, skewed results — describe a model that got worse. A backdoor does not make the model worse. NIST defines backdoor poisoning as causing a model "to perform an adversary-selected behaviour in response to inputs that follow a particular backdoor pattern", which means normal behaviour everywhere else. Those signs are worth watching; they cover one objective out of several, and a list that does not say which one leaves the reader believing they are covered.

Is there a detection method built specifically for backdoor poisoning?

There are several, and NIST records a limitation for each. The classes NIST discusses for backdoor poisoning include data sanitization, trigger reconstruction, and model inspection and sanitization. Data sanitization "achieves better results when the poisoning attack controls a relatively large fraction of training data but is not as effective against stealthy poisoning attacks", and Severi and colleagues found two standard methods of that kind, spectral signatures and Activation Clustering, "ineffective against clean-label backdoor poisoning on malware classifiers". Trigger reconstruction assumes the trigger "is at a fixed location in the poisoned training samples", and NIST warns that newer "semantic and functional backdoor triggers would also pose challenges to approaches based on trigger reconstruction or model inspection". NIST's own conclusion is that "Additional research is required to design strong backdoor mitigation strategies".

Why not just search the training data for poisoned samples?

It is a reasonable route and it is also NIST's own recommendation in part, alongside verifying downloads against published cryptographic hashes. NIST's own assessment of the filtering half is blunt, though: "detecting poisoned data within a large training corpus may be very difficult." The difficulty scales with the corpus, and modern training corpora are the largest artifacts in the pipeline.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. MITRE ATLAS, AML.T0020 Training Data Poisoning (collection 2026.08)MITRE
  3. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingarXiv, 10 Jan 2024
  4. Data Poisoning: The Exploitation of Generative AICrowdStrike, 19 Mar 2024

Guides that use this term