What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Side-channel model extraction

Side-channel model extraction is an attack that recovers a machine learning model's architecture or weights by observing the hardware running it, such as its timing, electromagnetic emanations or memory, with or without the model's answers. Because the attacker is not limited to inputs and outputs, a proven worst-case limit on extraction no longer applies.

Last reviewed

Key points

  • Side-channel model extraction reads a model off the hardware that runs it, with or without its answers. NIST lists side channels as the third extraction technique, beside computing weights algebraically and learning-based methods.
  • Jagielski and colleagues proved that some networks need an enormous number of queries to copy so that every output matches, but the proof assumes the attacker sees only inputs and outputs. If that assumption is broken, they note, such a copy becomes practical.
  • Batina and colleagues recovered a network's layers, neurons, activation functions and weights from a microcontroller's electromagnetic emanations and timing while feeding it chosen inputs.
  • Rakin and colleagues used Rowhammer from an unprivileged process on the same machine to leak weight bits from memory, then trained a substitute model on those bits and a small share of the training data. Their attack assumed the architecture was already known.
  • NIST warns that limiting queries, detecting suspicious queries and building side-channel-resistant architectures can all be circumvented by motivated and well-resourced attackers.

Side-channel model extraction is model extraction that reads the machine running a model, with or without its answers. NIST defines a side channel as inferring a secret from how a program runs, such as its timing or memory, or from its hardware’s power use and electromagnetic emanations.

How it leaks

Batina and colleagues (USENIX Security 2019) fed chosen inputs to neural networks on two microcontrollers while a probe recorded electromagnetic emanations and timing. From those traces they recovered the activation functions, the number of layers and neurons, the number of output classes, and the weights. They opened the chip packages, but say the findings hold without that.

Rakin and colleagues (DeepSteal, IEEE S&P 2022) needed no probe. An unprivileged process on the same machine as the victim used Rowhammer to read weight bits from memory. After 4,000 rounds it had the highest-order bit of about 90 percent of the weights, about 11.5 percent of all bits. A substitute model built from those bits and about 8 percent of the CIFAR-10 training data reached up to 89.42 percent accuracy against a ResNet-34 victim scoring 93.11 percent. DeepSteal assumed the architecture was already known.

Why the hardness proof does not cover it

Jagielski and colleagues prove that a specially built family of two-layer networks needs a number of queries that grows exponentially with its width to copy exactly, meaning a copy that matches the original’s output on every input. They then add that the result “assumes the adversary can only observe the input-output behavior” of the model being copied. “If this assumption is broken then functionally equivalent extraction becomes practical.” A side channel gives the attacker something besides inputs and outputs to observe.

Where definitions disagree

OWASP’s LLM10:2025 lists side-channel attacks beside API extraction, but describes attackers who “exploit input filtering techniques of the LLM” to harvest weights and architecture. That is a route through the application, and OWASP’s matching control is sandboxing the model’s network access. NIST and the papers mean physical and memory-level leakage. The two share a name and a goal, not a mechanism.

Questions and answers

How is side-channel model extraction different from ordinary model extraction?

Ordinary model extraction works only from what the model's API returns, either computing weights from the answers or training a copy on them. Side-channel model extraction reads information from the hardware running the model, such as timing, electromagnetic emanations or memory contents, sometimes alongside the answers and sometimes without them. The DeepSteal attacker saw no model outputs at all. NIST lists side channels as a separate extraction technique from algebraic and learning-based methods.

Does side-channel model extraction need physical access?

Not always. The CSI NN attack by Batina and colleagues assumed physical access to a microcontroller, noting that this can be remote via electromagnetic signals, and measured it with a nearby probe. The DeepSteal attack by Rakin and colleagues needed only an unprivileged process running on the same machine as the victim model, which leaked weight bits from memory using Rowhammer.

Do API rate limits stop side-channel model extraction?

Not reliably. NIST lists limiting user queries, and building architectures that resist side-channel attacks, among model extraction mitigations, and warns that motivated and well-resourced attackers can circumvent them.

What can an attacker do with a model stolen through a side channel?

Use it to attack the original. In DeepSteal, adversarial examples crafted on the substitute model transferred to the victim, coming within 25 percent of a white-box attack for ResNet models, where white-box means an attacker who holds the real weights. NIST notes that model extraction is often a step toward such attacks rather than an end goal.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. High Accuracy and High Fidelity Extraction of Neural Networks (later published at USENIX Security)arXiv, 3 Sep 2019
  3. CSI NN: Reverse Engineering of Neural Network Architectures Through Electromagnetic Side Channel (28th USENIX Security Symposium)USENIX, Aug 2019
  4. DeepSteal: Advanced Model Extractions Leveraging Efficient Weight Stealing in Memories (later published at IEEE S&P 2022)arXiv, 8 Nov 2021
  5. LLM10:2025 Unbounded Consumption, OWASP Top 10 for LLM ApplicationsOWASP