Definition · AI security
Fast gradient sign method
The fast gradient sign method (FGSM) is a one-step attack that builds an adversarial example by computing the gradient of a model's loss with respect to its input, then moving every input value by the same small amount in the direction of that gradient's sign. Goodfellow, Shlens and Szegedy introduced it in 2014.
Last reviewed
Key points
- The fast gradient sign method (FGSM) makes an adversarial example in one step, moving every input value by the same small amount in the direction that raises the model's loss.
- Goodfellow, Shlens and Szegedy introduced it in 2014 to support their argument that neural networks are fooled because they behave too linearly, not because they are too nonlinear.
- One gradient computation is all it costs, which is what made adversarial training on generated examples practical.
- Its speed is also its weakness. Iterative attacks built from FGSM steps are stronger, and models trained only against FGSM can learn to beat FGSM without becoming robust.
The fast gradient sign method, FGSM, builds an adversarial example in a single step. It asks the model one question: for each input value, which direction would raise the error? Then it moves every value a fixed small amount that way. Goodfellow, Shlens and Szegedy named it the “fast gradient sign method” in 2014, and NIST describes it as an efficient method for generating adversarial examples for deep learning.
How it works
FGSM takes one step. It computes the gradient of the model’s loss with respect to the input with one pass of backpropagation, the same calculation used to train the network. It keeps only the sign of each entry, plus or minus, and multiplies by epsilon, the largest change allowed per value. That perturbation is added to the input.
Goodfellow and colleagues built the method from an argument about why adversarial examples exist. They wrote that the primary cause is the networks’ “linear nature”. Nudge thousands of pixels each by an invisible amount, all in the helpful direction, and “we can make many infinitesimal changes to the input that add up to one large change to the output.”
The analogy stops at the single look. FGSM reads the direction once. A model whose response curves near the input can make that one reading wrong, which is why the attack is fast rather than strong.
Why it matters
FGSM showed how little it takes. On MNIST, with each pixel allowed to move by 0.25, it drove a simple one-layer classifier to a 99.9% error rate. On ImageNet, a change the size of the smallest step in an 8-bit image turned GoogLeNet’s 57.7% “panda” into a 99.3% “gibbon”.
Being cheap is what made adversarial training workable. Earlier examples came from a slow optimiser; Goodfellow and colleagues called FGSM a fast method “that makes adversarial training practical”, and used it to cut a multi-layer network’s error on FGSM examples from 89.4% to 17.9%.
Trade-offs
A single step is weak against a careful defender. Kurakin, Goodfellow and Bengio repeated it in small clipped steps to make a stronger iterative attack, and Madry and colleagues treat projected gradient descent as that multi-step variant. Besides speed, FGSM kept two advantages in Kurakin and colleagues’ ImageNet tests. Its examples moved between models better, so they suited black-box attacks, and they survived printing and photographing better than those from the iterative methods.
Training only against FGSM can teach the wrong lesson. Kurakin and colleagues found it gave no robustness to iterative attacks, and that FGSM leaks the true label, so they advised against using FGSM “to evaluate robustness to adversarial examples”. Tramèr and colleagues showed why: the model bends its loss surface until FGSM’s one reading misleads. “The model thus learns to generate weak perturbations, rather than defend against strong ones.”
The verdict was reopened in 2020. Wong, Rice and Kolter found that starting FGSM from a random point inside the allowed range made FGSM training “as effective as PGD-based training”, reaching 45% robust accuracy on CIFAR10 in 6 minutes on a single GPU. They named the failure that may have sunk earlier attempts “catastrophic overfitting”, in which robust accuracy against PGD on the training data suddenly drops to 0%.
Questions and answers
Why does FGSM use only the sign of the gradient?
Because the attack is limited by how far any single input value may move, not by the total size of the change. Under that limit, Goodfellow and colleagues show the change that most increases a linear model's output moves every value the full allowed amount in the direction of the gradient's sign. The sign keeps the direction and throws away the magnitude the limit makes irrelevant.
Is FGSM still used?
Yes, mostly as a cheap baseline and inside training. Kurakin and colleagues found single-step attacks transfer better between models than iterative ones, making them the best choice for black-box attacks in their tests, and Wong, Rice and Kolter showed in 2020 that FGSM training with a random starting point can match PGD training at a fraction of the cost. As a test of whether a model is robust, it is too weak on its own.
What is the difference between FGSM and PGD?
FGSM takes one step of full size. Projected gradient descent takes many small steps of the same sign-of-gradient kind, pulling the input back inside the allowed range after each one. Madry and colleagues describe FGSM as a one-step scheme for the same problem PGD solves in many steps, and PGD is the stronger attack.
Sources
- Explaining and Harnessing Adversarial ExamplesarXiv (ICLR 2015), 20 Dec 2014
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- Adversarial examples in the physical worldarXiv (ICLR 2017 workshop), 8 Jul 2016
- Adversarial Machine Learning at ScalearXiv (ICLR 2017), 4 Nov 2016
- Ensemble Adversarial Training: Attacks and DefensesarXiv (ICLR 2018), 19 May 2017
- Towards Deep Learning Models Resistant to Adversarial AttacksarXiv (ICLR 2018), 19 Jun 2017
- Fast is better than free: Revisiting adversarial trainingarXiv (ICLR 2020), 12 Jan 2020