Definition · AI security
Model poisoning
Model poisoning is an attack that corrupts a model by tampering with the data or weights it learns from, before the model is ever deployed. Because the flaw is trained in rather than introduced at runtime, the resulting model carries it everywhere it runs and can still pass ordinary evaluation.
Last reviewed
Key points
- Model poisoning corrupts a model before it is deployed, so the flaw ships inside the weights and travels to every place the model runs.
- What separates it from data poisoning is capability. Model poisoning needs control of the model's parameters; data poisoning needs only control of training data.
- NIST names two routes to that control, inserting a Trojan trigger into the model and sending malicious updates in federated learning. The EU AI Act adds a third, a pre-trained component that arrived already poisoned.
- A poisoned model passes its tests. BadNets found backdoored models perform normally on the defender's validation set and fail only when the attacker's trigger appears.
- Model poisoning is used narrowly for parameter control and broadly for the whole family of attacks that corrupt a model before deployment, so check which sense a document means before mapping a control to it.
Model poisoning corrupts a model before anyone deploys it. The attacker works on the supply chain: the training run, the update stream, the published checkpoint. The flaw ships inside the weights, and the attack is finished before the defender serves a first request — unlike prompt injection, which must be delivered to a running model every time.
A poisoned model passes its tests. Gu and colleagues backdoored a US traffic sign classifier in 2017 and found such models “exhibit equivalent performance on the defender’s validation sets, but can then be forced to fail in the field when a backdoor-triggering input is seen.” Retrained to read Swedish signs, it stayed backdoored. Nor is the artefact a reliable fallback: NIST states that organizations and researchers “may not be able to audit and identify vulnerabilities encoded into a model’s weights” the way open-source software can be audited. It names mechanistic interpretability as a proposed approach, not an available one.
How a model gets poisoned
Model poisoning takes control of the model’s parameters. NIST names two routes to that control: “generating a Trojan trigger and inserting it in the model”, or “sending malicious local model updates in federated learning”. The EU AI Act names a third — a pre-trained component that arrived already poisoned.
Every route is someone else’s write access to your model. Federated learning grants it to each client by design; a supply chain grants it to whoever built the checkpoint, and NIST calls robustness against supply-chain model poisoning “a critical open problem”. Data poisoning reaches the model through the training set instead, needing no access to the weights.
Where definitions disagree
“Model poisoning” is used both narrowly and broadly. NIST’s sense is narrow and mechanical: a poisoning attack “which operates through model control”. The broad sense, the one this page uses, is any attack that corrupts a model before deployment, however the tampering got in. Article 15(5) of the EU AI Act sits between them, naming data poisoning and model poisoning side by side as separate categories for a high-risk system.
“Data poisoning” now also names a defence. Tools that perturb images so that models trained on them learn the wrong thing were built for artists protecting work from unlicensed training, and NIST warns they “may become harmful in the hands of people with malicious intent.” The technique carries no intent, and neither does the word.
Questions and answers
Is model poisoning the same as data poisoning?
Not quite, and the difference is what the attacker has to control. NIST defines model poisoning as a poisoning attack that "operates through model control", meaning control of the model's parameters, and data poisoning as one where "an adversary controls a subset of the training data". Both act during training, so both are poisoning, but they need different access and different controls stop them. In Article 15(5) of the EU AI Act, "model poisoning" also covers a poisoned pre-trained component inherited from a supplier. This page uses that broad sense and names the narrow one where it changes which defence applies.
Can you tell a poisoned model by testing it?
Not by ordinary evaluation. BadNets showed in 2017 that backdoored models "exhibit equivalent performance on the defender's validation sets, but can then be forced to fail in the field when a backdoor-triggering input is seen", so a team that evaluated the model the way it evaluates every model would ship it. Inspecting the artefact is not a reliable fallback either. NIST states that organizations and researchers "may not be able to audit and identify vulnerabilities encoded into a model's weights in the same way it is often possible to audit open-source software". That is a limit rather than a closed door, and NIST says so about model poisoning by name: new approaches are required to detect vulnerabilities in models such as those introduced through model poisoning attacks, and it records mechanistic interpretability, used to identify backdoor features, among the approaches currently proposed. Proposed is not the same as available.
Does federated learning make model poisoning more likely?
It is one of the two places NIST says the attack is most prevalent. In federated learning "clients send local model updates to the aggregating server", which hands every client a legitimate channel for writing to the model's parameters — exactly the capability model poisoning requires. The other place is the supply chain, "where malicious code may be added to the model by suppliers of model technology". Both grant parameter access by design, which is why neither is fixed by controlling what enters the training set.
Sources
- BadNets, Identifying Vulnerabilities in the Machine Learning Model Supply ChainarXiv, 22 Aug 2017
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, Robustness and CybersecurityEuropean Union, 12 Jul 2024