What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Architectural backdoor

An architectural backdoor is a backdoor built into a machine learning model's structure, the layers and operations set out in the model's code, rather than into its weights or training data. It makes the model misbehave when an input carries the attacker's trigger, and some designs survive retraining because they do not depend on learned weights.

Last reviewed

Key points

  • An architectural backdoor lives in a model's structure, the layers and operations the model's code defines, not in its weights or its training data.
  • Some designs survive retraining because they do not depend on learned weights. Bober-Irizar and colleagues showed one surviving "a complete re-training from scratch", and found that designs lacking their requirements struggle to survive.
  • The attacker needs access to the model design or the training environment, which NIST says applies when training is outsourced, for example to a cloud service.
  • A simple route is to publish a backdoored model definition for others to reuse. In a small study, ML practitioners mostly failed to find these backdoors, even after being warned they might be present.

How it works

A backdoor makes a model misbehave only when an input carries the attacker’s trigger. A model has two parts: the structure its code defines, and the weights training fills in. Model poisoning and data poisoning reach the backdoor through the weights. An architectural backdoor is written into the structure.

Langford and colleagues describe building one from “common architectural components, such as activation functions and pooling layers”. The added components detect the trigger and change the output when it appears. In a minimal implementation, their backdoors added between 6 and more than 100 lines to the model’s code, with a median of 9.

Survival depends on the design. Bober-Irizar and colleagues found that without certain requirements, such as a weight-agnostic implementation, “the learned backdoors will struggle to survive re-training”. Langford and colleagues built trigger detectors that work for any weights, and state they “always provably survive any amount of retraining or fine-tuning”.

Why it matters

The attacker does not need your data. NIST says the attack requires “adversarial access to the model design or training environment” and applies when training is outsourced, for example to a cloud service.

Bober-Irizar and colleagues name a second route: “publishing open-source code for a backdoored model architecture that others will reuse unknowingly”. Langford and colleagues note that users commonly load model definitions from hubs or GitHub unchanged, which makes this an AI supply chain compromise. They list full retraining from scratch among common practices for removing backdoors, and a weight-agnostic design survives it.

Reading the code is a weak check. In their study of 11 practitioners, warned that backdoors might be present, “subjects generally struggled to identify backdoors”.

Where definitions disagree

NIST describes the modification as happening “during its training phase”. MITRE ATLAS’s Modify AI Model Architecture technique is not tied to training. One of its case studies, a research exercise rather than a real incident, injected a payload into compiled Android app models “by directly modifying the computation graph”, the operations the saved model runs. Langford and colleagues, the paper NIST cites for its passage, say backdoors can be injected “both before and after training”. NIST’s wording describes the training-time case; ATLAS also covers edits to a finished model.

Questions and answers

Does retraining a model remove an architectural backdoor?

It depends on how the backdoor is built. Bober-Irizar and colleagues showed an architectural backdoor that can "survive a complete re-training from scratch", but found that without their construction requirements "the learned backdoors will struggle to survive re-training". Langford and colleagues state that architectural backdoors "can be injected in such a way that they always provably survive any amount of retraining or fine-tuning", because their constructions work for any weights. Retraining replaces the weights, so it cannot remove a backdoor that does not depend on them.

How is an architectural backdoor different from model poisoning?

Model poisoning changes a model's weights; an architectural backdoor changes its structure. MITRE ATLAS files both as sub-techniques of Manipulate AI Model: Poison AI Model covers the weights and Modify AI Model Architecture covers layers and operations. The practical difference is that retraining replaces the weights but leaves the structure, so a structural backdoor built not to depend on weights stays in place.

Can code review catch an architectural backdoor?

Not reliably, on the evidence so far. Langford and colleagues gave 11 ML practitioners two minutes per model, told them the architectures may contain backdoors, and found that "subjects generally struggled to identify backdoors", with most guesses random. The six backdoors were "specifically engineered to be challenging to detect" and sat in network definitions that typically span hundreds or thousands of lines. The authors say the sample is too small for statistical inference.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. Architectural Backdoors in Neural NetworksarXiv, 15 Jun 2022
  3. Architectural Neural Backdoors from First PrinciplesarXiv, 10 Feb 2024
  4. MITRE ATLAS, AML.T0018.001 Modify AI Model Architecture (collection 2026.09)MITRE
  5. DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload InjectionarXiv (ICSE 2021), 18 Jan 2021