What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Backdoor attack

A backdoor attack builds a hidden conditional rule into a model during training, so the model behaves normally on ordinary inputs and produces the attacker's chosen behaviour on any input carrying a particular trigger. The backdoor is the rule left in the model; poisoning the training data is one of several routes to it.

Last reviewed

Key points

  • A backdoor is a conditional rule in the trained model — ordinary behaviour on ordinary inputs, the attacker's chosen behaviour on any input carrying their trigger.
  • Backdoor and poisoning are not the same thing. Poisoning the training data is one route to installing the rule; NIST says model poisoning and architectural changes are others.
  • Installing the backdoor is only half the attack. Within its poisoning taxonomy, NIST says backdoor poisoning is the only attack needing both training and testing data control — the poison goes in, and someone must present the trigger later.
  • The count matters more than the share. Souly and colleagues backdoored models from 600M to 13B parameters with 250 poisoned documents, 0.00016 percent of tokens for the largest.
  • Survival is cheap; a large gap is bought. Retrained for Swedish signs, a backdoored US sign classifier read clean images at 74.9 percent and triggered ones at 61.6; BadNets' 25 percent needed a strengthened backdoor.

NIST’s example of a backdoor trigger is as plain as “the insertion of a square of white pixels”. Gu and colleagues built the canonical case in 2017, a street sign classifier that reads stop signs as speed limits once a sticker is on them.

The backdoor is the property, poisoning is the route

“Backdoor” and “poisoning” get used as synonyms and name different things. NIST names routes that never touch the training set: model poisoning “could result in either targeted or backdoor attacks”, reaching the parameters directly, while architectural backdoors modify the model’s structure during training.

Installing the rule is only half the attack. Inside its poisoning taxonomy, NIST calls backdoor poisoning “the only attacks in the literature that require both training and testing data control”: someone must present the trigger afterwards. MITRE ATLAS splits the halves into separate techniques: Poison AI Model installs it, Insert Backdoor Trigger fires it, “in conjunction with” the first.

How few samples it takes

Fewer than a percentage of the corpus suggests, and the number does not grow with the model. Souly and colleagues found in 2025 that “250 poisoned documents similarly compromise models across all model and dataset sizes”, 600M to 13B parameters — “only 0.00016% of training tokens” at 13B. At 100 documents they saw no successful poisoning.

Their scope was narrow and they say so: gibberish output and language switching, and they “do not demonstrate any successful end-to-end poisoning attacks”. Survival is cheap; a large gap is bought. BadNets’ backdoor came through retraining for Swedish signs unamplified, and only strengthening it bought the headline 25 percent gap; NIST records latent backdoors engineered for it, “designed to survive even upon model fine-tuning of the last few layers using clean data”.

Where definitions disagree

“Trojan” names both halves. NIST’s glossary calls it “a malicious modification to a model”, the implant, then says “trojan may be used interchangeably with backdoor pattern”, the trigger. One word for the lock and the key. ATLAS goes further and has no backdoor technique at all, listing embedded backdoors among the effects of training data poisoning.

Questions and answers

What is the difference between a backdoor and data poisoning?

A backdoor is a property of the trained model; poisoning is a method of putting it there. The distinction holds because the routes are not the same. NIST states that "model poisoning attacks could result in either targeted or backdoor attacks", where the attacker modifies model parameters and never touches the training set, and describes architectural backdoor attacks that modify "the structure of an ML model during its training phase". The terms are used interchangeably in practice and are not interchangeable in fact.

How many poisoned samples does it take to install a backdoor?

Far fewer than a percentage of the corpus implies. Souly and colleagues pretrained models from 600M to 13B parameters in 2025 and found that "250 poisoned documents similarly compromise models across all model and dataset sizes" — for the 13B model, "only 0.00016% of training tokens". The absolute count mattered and the share did not, so the attack does not get harder as models get bigger. Read the scope with it, though. The authors tested a narrow set of backdoors, gibberish output and language switching, and state they "do not demonstrate any successful end-to-end poisoning attacks" because they did not show persistence through realistic post-training.

Does fine-tuning remove a backdoor?

Not dependably, and the good backdoors are built not to be removed that way. Gu and colleagues retrained a backdoored US street sign classifier for Swedish signs. It scored 74.9 percent on clean images — 2.2 points above an honestly trained Swedish baseline — and 61.6 percent once the trigger was present. The 25 percent figure the paper leads with belongs to a further experiment that amplifies the backdoor neurons: at k = 20, "in return for a 3% drop in accuracy for clean images, this attack causes a > 25% drop in accuracy for backdoored images". Survival is cheap; a large gap is bought. NIST records latent backdoor attacks as "designed to survive even upon model fine-tuning of the last few layers using clean data" — designed to, which is the point. Survival is an engineering goal of the attack rather than a property of backdoors in general. Souly and colleagues, working on language models, found the opposite tendency for their attacks, that "continued clean training may eventually remove backdoors in certain settings".

Is a backdoor trigger the same as an adversarial example?

No, and the difference is whether anyone tampered with the model. An adversarial example exploits a weakness already present in an honestly trained model. A backdoor trigger fires a rule an attacker installed during training, which is why NIST says backdoor poisoning attacks "are the only attacks in the literature that require both training and testing data control". An adversarial example needs only the second.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. BadNets, Identifying Vulnerabilities in the Machine Learning Model Supply ChainarXiv, 22 Aug 2017
  3. Poisoning Attacks on LLMs Require a Near-constant Number of Poison SamplesarXiv (UK AI Security Institute, Anthropic, Alan Turing Institute), 8 Oct 2025
  4. MITRE ATLAS, AML.T0043.004 Insert Backdoor Trigger and AML.T0018.000 Poison AI Model (collection 2026.08)MITRE

Guides that use this term

In the news