Definition · AI security
Gradient masking
Gradient masking is a failure mode of defences against adversarial examples: the defence leaves the model without useful gradients, so attacks that follow the gradient to find an adversarial input fail. The adversarial inputs are still there. Attacks that do not rely on the defended model's own gradient, such as transfer from a substitute model, can often still find them.
Last reviewed
Key points
- Gradient masking is when a defence leaves a model without useful gradients, so attacks that follow the gradient to find an adversarial example stop working.
- The adversarial examples have not gone away. Carlini and colleagues say masking may make optimisation-based attacks fail, but that does not mean the space of adversarial perturbations shrank.
- Papernot and colleagues showed that masked models can still be beaten by attacks built on a separate substitute model and transferred.
- Athalye, Carlini and Wagner found that 7 of 9 defences at ICLR 2018 that claimed white-box robustness without a proof relied on obfuscated gradients, a special case of gradient masking, and fully circumvented 6 of them.
- One warning sign is a black-box attack doing better than a white-box one. A black-box attacker has strictly less to work with, so without masking it should do worse.
A gradient tells an attacker which way to nudge an input to raise a model’s error. Attacks such as the fast gradient sign method and projected gradient descent follow it to build an adversarial example. Gradient masking breaks that guide instead of fixing the model.
How it works
Papernot and colleagues named the pattern: defences that construct “a model that does not have useful gradients”. There are several ways to get there.
- Break the maths. A step that cannot be differentiated, such as discretising the input, leaves no usable gradient. Athalye, Carlini and Wagner call these shattered gradients.
- Add randomness. Transforming each input at random makes a single gradient sample misleading.
- Go very deep. Looping the input through a network many times can make gradients vanish or explode.
- Learn it by accident. Tramèr and colleagues found that adversarial training with a weak one-step attack curves the loss sharply near each input, so the gradient “poorly reflects the global loss landscape”.
The first three are designed in. Athalye, Carlini and Wagner call those obfuscated gradients, “a special case of gradient masking”.
Why it matters
A masked model can pass an evaluation that only runs gradient attacks, and that result proves little. Carlini and colleagues put it plainly: gradient masking “may make optimization-based attacks fail but that does not mean that the space of adversarial perturbations decreased”.
One way around it is to take a gradient from somewhere else. Papernot and colleagues crafted attacks on a separate substitute model. Defensive distillation stopped the fast gradient sign method run on the model itself, but the same attack through a substitute succeeded “regardless of the distillation temperature”. This is transferability at work.
In practice
Athalye, Carlini and Wagner rebuilt the gradient instead: they swapped a broken step for a smooth stand-in on the backward pass (BPDA) and averaged gradients over the randomness. They studied the defences at ICLR 2018 that claimed robustness against white-box attacks without a proof. Of those 9, 7 relied on obfuscated gradients. They circumvented 6 fully and 1 partially.
Thermometer encoding is a clear case. It discretises each pixel so gradient descent cannot run directly. Against its authors’ own attack, encoding alone kept 50% accuracy on CIFAR-10 images when no pixel could change by more than 0.031 on a 0-to-1 scale. Using BPDA, Athalye, Carlini and Wagner cut that to 0% at the same limit, “no measurable improvement” over an undefended model.
The test for masking is to run attacks that ought to be weaker. White-box attacks can do everything black-box ones can, so they should win. Carlini and colleagues call a white-box attack doing worse than a gradient-free one “strong evidence for gradient masking”. Athalye, Carlini and Wagner list more signs: one-step attacks beating iterative ones, random search finding examples that gradient attacks miss, and success that does not grow with the allowed distortion.
Questions and answers
Is gradient masking the same as obfuscated gradients?
Not quite. Athalye, Carlini and Wagner define obfuscated gradients as "a special case of gradient masking": the case where a defence is designed in a way that necessarily breaks the gradient, such as a non-differentiable preprocessing step or test-time randomness. Gradient masking also covers models that learn bad gradients during training, such as adversarial training with weak single-step attacks.
How do you tell whether a defence is masking gradients?
Compare attacks that should be weaker with ones that should be stronger. Carlini and colleagues call a white-box attack doing worse than a gradient-free attack "strong evidence for gradient masking". Athalye, Carlini and Wagner add other warning signs: one-step attacks beating iterative ones, random search finding examples gradient attacks miss, and success that does not rise as the allowed distortion grows.
Sources
- Practical Black-Box Attacks against Machine LearningarXiv (Papernot et al.), 19 Mar 2017
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial ExamplesarXiv (Athalye et al.), 31 Jul 2018
- On Evaluating Adversarial RobustnessarXiv (Carlini et al.), 20 Feb 2019
- Ensemble Adversarial Training: Attacks and DefensesarXiv (ICLR 2018), 19 May 2017