What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Defensive distillation

Defensive distillation is a defence against adversarial examples, proposed by Papernot and colleagues in 2016, that trains a classifier on the softened probability outputs of an identical classifier trained at a high softmax temperature. It stopped the attacks of its day, but later work showed it hides the gradients those attacks follow rather than making the model robust.

Last reviewed

Key points

  • Defensive distillation trains a classifier at a high softmax temperature, then trains a second network of the same size on its softened probabilities and runs that one at temperature 1.
  • Papernot and colleagues reported in 2016 that, at temperature 100, it cut their attack's success from 95.89% to 0.45% on MNIST and from 87.89% to 5.11% on CIFAR-10, with the attack capped at 112 changed input features.
  • Carlini and Wagner showed that the distilled model's gradients usually round to zero, and built attacks that found adversarial examples for 100% of images on distilled networks.
  • Papernot and colleagues found that the fast gradient sign method, run through an undistilled substitute model, beat distilled MNIST models at every temperature they tried (5, 10 and 100).
  • Papernot and colleagues later classed it as gradient masking. It is not model distillation, which makes a smaller or cheaper model.

A classifier’s softmax layer turns raw scores into probabilities. Its temperature setting controls how spread out they are. Defensive distillation turns it up to resist adversarial examples.

How it works

Papernot and colleagues proposed three steps.

  1. Train a classifier with the softmax at a high temperature, such as 100.
  2. Train a second classifier of the same size on its soft probabilities, not the right answers, also at that temperature.
  3. Run the second classifier at temperature 1.

The method comes from model distillation, which uses the same training to make a smaller copy.

With the distilled networks trained at temperature 100, success of the authors’ attack (capped at 112 changed input features) fell from 95.89% to 0.45% on a network for handwritten digits (MNIST) and from 87.89% to 5.11% on one for small photos (CIFAR-10).

Why it matters

Defensive distillation passed its own evaluation without making the model robust. Papernot and colleagues later classed it as gradient masking.

Carlini and Wagner explained why some attacks failed. Training at 100 and running at 1 makes the model’s raw scores about 100 times larger. The output becomes almost exactly 1 for one class and 0 for the rest. The gradient, which tells an attacker which way to nudge the input, usually becomes so small that it rounds to 0 in 32-bit arithmetic, so attacks that follow it make no progress. They could not fully explain why others, such as the fast gradient sign method, failed.

The adversarial examples were still there. Carlini and Wagner built attacks that found them for 100% of images on distilled networks. On MNIST, the average distance from an image to its adversarial example showed no effect of temperature from 1 to 100. A higher temperature, they concluded, “only causes existing attacks to fail more often”.

In practice

Defensive distillation was broken twice more, by taking the gradient from somewhere else. A team led by its first author, Papernot, trained an undistilled substitute model and attacked that. The attack then worked on the distilled model “regardless of the distillation temperature”, at each temperature they tried (5, 10 and 100). They suggest this is because the substitute, never distilled, has the gradients the attack needs. Carlini and Wagner showed that high-confidence adversarial examples from an ordinary model also transfer to a distilled one. Both are cases of transferability.

Carlini and Wagner said the case “illustrates the general need for better techniques to evaluate the robustness of neural networks”.

MITRE ATLAS still lists defensive distillation among the robustness techniques that “may” make up its mitigation AML.M0003, Predictive AI Model Hardening, alongside adversarial training. The listing does not say how well each technique works.

Questions and answers

Is defensive distillation the same as model distillation?

No. Both train one model on another model's probability outputs, but model distillation makes a smaller or cheaper copy. Carlini and Wagner point out that in defensive distillation the two models are "identical in size", and the high temperature is there to make the final model more confident, not smaller.

Does defensive distillation work?

Not as a defence against an attacker who adapts. It stopped the attacks its authors tested, but Carlini and Wagner found adversarial examples for 100% of images on distilled networks, and Papernot and colleagues beat it with a substitute model at every temperature they tried.

Sources

  1. Distillation as a Defense to Adversarial Perturbations against Deep Neural NetworksarXiv (Papernot et al., IEEE S&P 2016), 14 Mar 2016
  2. Towards Evaluating the Robustness of Neural NetworksarXiv (Carlini and Wagner, IEEE S&P 2017), 22 Mar 2017
  3. Practical Black-Box Attacks against Machine LearningarXiv (Papernot et al.), 19 Mar 2017
  4. MITRE ATLAS, AML.M0003 Predictive AI Model Hardening (collection 2026.09)MITRE