What matters in AI.

Subscribe

Learn / AI basics

Definition · AI basics

Model distillation

Model distillation is a training technique in which a smaller student model learns to reproduce the behaviour of a larger teacher model by training on the teacher's outputs, often its full probabilities over possible answers, instead of on correct labels alone. The goal is a cheaper model that keeps most of the teacher's capability.

Last reviewed

Key points

  • Model distillation trains a small student model on the outputs of a larger teacher model, so the student learns how the teacher answers rather than only what the right answer is.
  • The classic form trains on the teacher's probability for every possible answer, which shows how wrong each wrong answer is. Training on text the teacher wrote is used as well.
  • Anthropic says frontier labs routinely distil their own models to ship smaller, cheaper versions for their customers.
  • The technique carries no intent. When three labs distilled Claude through fraudulent accounts, Anthropic called it illicit distillation and MITRE ATLAS filed the step as model extraction.

Model distillation copies what a large model has learned into a small one. Hinton, Vinyals and Dean named it in 2015 as a way “to transfer the knowledge from the cumbersome model to a small model that is more suitable for deployment”, crediting Rich Caruana and collaborators with showing earlier that it could be done. The large model is usually called the teacher and the small one the student.

How it works

The student learns from the teacher’s answers rather than from a bare answer key. A normal classifier is trained only to say that a picture shows a BMW. A teacher also gives a probability for every other answer, and the 2015 paper’s example is that a BMW “may only have a very small chance of being mistaken for a garbage truck, but that mistake is still many times more probable than mistaking it for a carrot.” Those relative odds, called soft targets, show how the teacher generalises. The authors generally got their best results by also training on the correct labels, at a considerably lower weight.

Distillation also works on text the teacher wrote. Google’s Gemma 2 2B and 9B models, already trained on a larger teacher’s next-token probabilities, were tuned further on responses “predominantly synthetically generated by the teacher”.

Why it matters

Distillation lets a lab ship a model that is cheap to run. In the 2015 paper’s speech tests, one distilled model kept more than 80 percent of the gain in per-frame accuracy that came from averaging ten models, on a slightly outdated version of the speech model behind Android voice search. DistilBERT, from Hugging Face, made a copy of BERT 40 percent smaller that kept 97 percent of its score on GLUE, a standard set of language-understanding tests. Anthropic says frontier labs “routinely distill their own models to create smaller, cheaper versions for their customers.”

When distillation is an attack

Distillation turns into an attack when the teacher is another lab’s model, copied without permission. Anthropic reported in February 2026 that DeepSeek, Moonshot and MiniMax generated over 16 million exchanges with Claude through about 24,000 fraudulent accounts, “in violation of our terms of service and regional access restrictions”. The goal, Anthropic says, was either responses “for direct model training” or tasks “needed to run reinforcement learning”.

Anthropic still calls distillation “a widely used and legitimate training method” and names the abuse “illicit distillation”. MITRE ATLAS files the same step as Extract AI Model, its entry for model extraction. The training method does not give the attack away. Anthropic spotted it in the traffic: “What distinguishes a distillation attack from normal usage is the pattern.”

Questions and answers

What is model distillation?

Model distillation is a way of training a small model, the student, to behave like a larger one, the teacher. Instead of learning only from correct answers, the student learns from the teacher's outputs, often its probability for every possible answer. The result is a model that is cheaper to run and keeps most of the teacher's capability.

Who invented knowledge distillation?

Hinton, Vinyals and Dean named it "distillation" in a 2015 paper, but they credit Rich Caruana and his collaborators with first showing that the knowledge in an ensemble of models can be compressed into a single model. The 2015 paper contributed a more general method, which softens the teacher's probabilities using a setting called temperature.

Is model distillation an attack?

No. Distillation is an ordinary training technique, and Anthropic calls it "a widely used and legitimate training method" that labs use on their own models. It becomes an attack when someone distils another lab's model without permission, as Anthropic reported three labs did to Claude in 2026. Anthropic calls that "illicit distillation", and MITRE ATLAS files it as model extraction.

Can a model be distilled through its API?

Yes, from its text responses alone. Anthropic reports that one goal of the labs distilling Claude was to collect "high-quality responses for direct model training". Legitimate distillation uses teacher-written text too, as when Google fine-tuned Gemma 2 on responses "predominantly synthetically generated by the teacher".

Sources

  1. Distilling the Knowledge in a Neural NetworkarXiv, 9 Mar 2015
  2. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighterarXiv, 2 Oct 2019
  3. Gemma 2: Improving Open Language Models at a Practical SizearXiv, 31 Jul 2024
  4. Detecting and preventing distillation attacksAnthropic, 23 Feb 2026
  5. MITRE ATLAS, AML.CS0056 and AML.T0024.002 Extract AI Model (collection 2026.09)MITRE