Definition · AI security
Projected gradient descent
Projected gradient descent (PGD) is an iterative attack that builds an adversarial example by repeatedly stepping the input in the direction that increases a model's loss, then projecting it back into the set of allowed perturbations, usually a small ball around the original. Croce and Hein call it the most popular way to test a model's robustness.
Last reviewed
Key points
- Projected gradient descent (PGD) builds an adversarial example in many small steps, each moving the input in the direction that raises the model's loss, then projecting it back inside the allowed range of change.
- Madry and colleagues conjectured in 2017 that no attack using only the model's gradients does much better, and made PGD the attack adversarial training trains against.
- When Athalye, Carlini and Wagner broke most non-certified ICLR 2018 defences, they judged the PGD-trained model free of the obfuscated gradients that undid most others.
- Croce and Hein showed that PGD's fixed step size and usual loss can overestimate robustness, and built Auto-PGD, which sets its own step size.
Projected gradient descent, PGD, builds an adversarial example in many small steps rather than one. Each step nudges the input the way that raises the model’s loss, a score of how wrong its answer is. Then the input is pulled back inside the allowed change, so the final result stays close to the original. NIST lists it among the notable white-box evasion attacks.
How it works
PGD starts from the original input, or from a random point near it. Each step computes the gradient, the direction in which the loss rises fastest, moves a small amount that way, then projects back into the allowed set. For the common per-pixel limit, projecting means clipping each value to within a fixed distance, epsilon, of the original.
In its common per-pixel form, each step is the fast gradient sign method (FGSM) at a smaller size. Kurakin, Goodfellow and Bengio first repeated it this way, as the “basic iterative” method. Madry and colleagues framed it as projected gradient descent, started it from a random point near the input, and in their MNIST attack tests restarted it up to 20 times.
Why it matters
For five MNIST and five CIFAR10 images, Madry and colleagues restarted PGD from 100,000 random points each and found the peaks it reached had similar loss values. They conjecture that an attack using only gradients “will not find significantly better local maxima than PGD”.
That made PGD the recipe for adversarial training. On CIFAR10, against 20-step PGD with a per-pixel limit of 8/255, a wide network trained on PGD examples kept 45.8% accuracy. Trained on normal images it kept 3.5%; trained on FGSM examples, 0.0%.
Athalye, Carlini and Wagner found 7 of 9 non-certified white-box defences at ICLR 2018 relied on obfuscated gradients, which hide the true slope from an attacker, and broke six in full. They believe the PGD-trained model does not, and report 47% accuracy for it on CIFAR at epsilon 0.031, against the per-pixel limit it was trained on.
Trade-offs
Every PGD step needs its own gradient computation, so a 20-step PGD attack costs about 20 times as much as one FGSM step. Wong, Rice and Kolter write that adversarial training often multiplies training time by an order of magnitude.
Croce and Hein name two reasons PGD can fail and overestimate robustness: its fixed step size and the cross-entropy loss it usually climbs. Their ensemble of attacks, run on over 50 published models, lowered the reported robust accuracy of all but one, often by more than 10%.
Where definitions disagree
The untargeted attack raises the loss, but the name says descent. Madry and colleagues mean descent “on the negative loss function”. NIST’s taxonomy says PGD “minimizes the loss function”, which fits that negative, or a targeted attack that lowers the loss for a chosen wrong answer. The authors of one 2018 defence named their variant “Logit-Space Projected Gradient Ascent”. The direction depends on which loss is written down.
Where the basic iterative method ends and PGD begins is also loose. Wong, Rice and Kolter trace PGD adversarial training to the basic iterative method, which “was further strengthened by adding multiple random restarts” and folded into training. Madry and colleagues name it after the general optimisation method, “the standard method for large-scale constrained optimization”, and add a random start that the basic iterative method lacked.
Questions and answers
What is the difference between PGD and FGSM?
The fast gradient sign method takes one full-size step in the direction that raises the model's loss. In the common per-pixel form, PGD takes many smaller steps of the same kind and pulls the input back inside the allowed range after each one. Madry and colleagues describe FGSM as a one-step scheme for the problem PGD solves in many steps, and PGD is the stronger attack.
Why is it called descent if the attack increases the loss?
Madry and colleagues describe it as projected gradient descent "on the negative loss function", which is the same as climbing the loss. The name comes from the optimisation method. At least one variant, built against a 2018 defence, was named Logit-Space Projected Gradient Ascent.
Does surviving PGD prove a model is robust?
No. Madry and colleagues conjecture that no gradient-based attack will do much better than PGD, but say their tests do not rule out rare, much stronger points. Croce and Hein showed that PGD's fixed step size and usual loss can make it overstate robustness, and their stronger ensemble lowered the reported robustness of nearly every model it tested.
Sources
- Towards Deep Learning Models Resistant to Adversarial AttacksarXiv (ICLR 2018), 19 Jun 2017
- Adversarial examples in the physical worldarXiv (ICLR 2017 workshop), 8 Jul 2016
- Fast is better than free: Revisiting adversarial trainingarXiv (ICLR 2020), 12 Jan 2020
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial ExamplesarXiv (ICML 2018), 1 Feb 2018
- Reliable Evaluation of Adversarial Robustness with an Ensemble of Diverse Parameter-free AttacksarXiv (ICML 2020), 3 Mar 2020