What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Model extraction

Model extraction is an attack that rebuilds a working copy of a machine learning model from its outputs alone, by repeatedly querying a prediction API and training a replacement on the answers it returns. Model extraction targets the model as property, rather than the private training data that other inference-time privacy attacks pursue.

Last reviewed

Key points

  • Model extraction rebuilds a working copy of a model from its answers alone. The attacker queries the API, keeps the outputs as labels, and trains a replacement.
  • The motive is economic. ATLAS records adversaries extracting models to avoid paying per query, and rates Extract AI Model Realized while the two privacy attacks filed beside it are only Feasible.
  • A stolen copy is a stepping stone rather than an end. NIST notes extraction is often a step toward other attacks, because a local copy turns a black-box target into a white-box one.
  • Distillation is the same operation performed with permission. What changes between the two words is authorisation, not technique.
  • Hardness results bound only the precise end of the problem. Exact extraction is fundamentally impossible for many classes of neural network; a copy that is merely as good at the task is not, and that is the copy a competitor wants.

Model extraction turns a model’s answers back into the model. MITRE ATLAS files it as AML.T0024.002: the adversary queries the victim’s inference API and uses the answers “as labels for training a separate model offline that will mimic the behavior and performance of the target model.” Every query is a small authorised disclosure, and enough of them add up to a working replacement. Membership inference and model inversion are filed beside it, but they target the training data; extraction targets the model.

Why a copy is worth building

The copy is cheap and the original was not. Anthropic describes the payoff as capabilities acquired “in a fraction of the time, and at a fraction of the cost, that it would take to develop them independently”, and reported in February 2026 that three labs generated “over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts”.

The copy is also a stepping stone. NIST notes that extraction “is often not an end goal but a step toward other attacks”: a copy the attacker runs locally is a white-box surrogate for a black-box target, which makes it easy to craft adversarial examples that transfer. It inherits capability without restraint, too. Anthropic reports that models built by illicit distillation are unlikely to retain the original’s safeguards, which is why ATLAS ends its case study on societal and user harm, not on the theft.

Where definitions disagree

A hardness result is not a defence. Jagielski and colleagues call exact extraction “purely a strawman”, and prove that checking whether two networks are functionally equivalent is NP-hard. Read as “you cannot steal the weights”, that is true and nearly irrelevant. No published hardness result bounds task accuracy extraction — “the easiest goal to consider in this taxonomy”, and the only one a competitor needs. Carlini and colleagues recovered one layer of OpenAI’s ada and babbage for under 20 US dollars; the gap between hard in theory and done for the price of the API calls is where the confusion lives. Separately, Jagielski and colleagues show that copying a specially built family of networks so every output matches, given raw scores, needs queries exponential in the network’s width. That worst case assumes an attacker who sees only inputs and outputs, so side-channel model extraction, which watches the hardware running the model, falls outside it.

Distillation is the same operation with permission, named by Hinton, Vinyals and Dean in 2015. ATLAS files the procedure twice, as Extract AI Model under Exfiltration and as AML.T0005.001 Train Proxy via Replication under staging, where the copy exists to craft adversarial data. OWASP files model theft under LLM10:2025 Unbounded Consumption; NIST files it under privacy.

Questions and answers

What is model extraction?

Model extraction is an attack in which an adversary rebuilds a usable copy of a machine learning model using only the answers its API returns. The attacker sends a large volume of queries, keeps the outputs as labels, and trains a replacement model on that dataset. No access to the original weights, training data or source code is required.

Is model extraction the same as distillation?

The operation is the same and the authorisation is not. Distillation was named by Hinton, Vinyals and Dean in 2015 — "a different kind of training, which we call 'distillation' to transfer the knowledge from the cumbersome model to a small model that is more suitable for deployment" — and labs distil their own models routinely. The term "model extraction" is used when the model being copied belongs to someone else and the copying breaches their terms. Nothing in the technique itself marks the difference.

Can an attacker steal the actual weights of a production model?

Only partially, and only under specific conditions. Carlini and colleagues recovered the embedding projection layer of OpenAI's ada and babbage models for under 20 US dollars in 2024, but the attack required an API that exposed full log probabilities or a logit bias, and both OpenAI and Google deployed defences after disclosure. Recovering a whole modern model's parameters from queries alone remains out of reach, which is why attackers target the behaviour instead.

Does differential privacy protect against model extraction?

No. NIST states plainly that differential privacy does not provide guarantees against model extraction, because it is designed to protect the training data rather than the model. Differential privacy bounds what an attacker learns about individual records, and model extraction is not trying to learn about individual records.

Do rate limits stop model extraction?

Rate limits raise the cost of model extraction without ending it. ATLAS notes that query limits may not protect against attacks that require few requests, and advises monitoring for evasion through distributed requests, account rotation or stolen credentials. The distillation campaigns Anthropic reported in February 2026 spread roughly 16 million exchanges across about 24,000 accounts, which is what a per-account limit looks like when the attacker is well funded.

Sources

  1. MITRE ATLAS, AML.T0024.002 Extract AI Model (collection 2026.08)MITRE
  2. Detecting and preventing distillation attacksAnthropic, 23 Feb 2026
  3. High Accuracy and High Fidelity Extraction of Neural Networks (later published at USENIX Security)arXiv, 3 Sep 2019
  4. Stealing Part of a Production Language Model (later published in Proceedings of the 41st International Conference on Machine Learning, PMLR 235)arXiv, 11 Mar 2024
  5. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  6. LLM10:2025 Unbounded Consumption, OWASP Top 10 for LLM ApplicationsOWASP
  7. Distilling the Knowledge in a Neural NetworkarXiv, 9 Mar 2015

Guides that use this term