What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Adversarial training

Adversarial training is a defence that trains a machine learning model on adversarial examples generated during training, each carrying its correct label, so the model learns to classify inputs an attacker deliberately perturbed. NIST names it one of three evasion defences that have proven resilient, and its resistance grows with the strength of the attack used to make the examples.

Last reviewed

Key points

  • Adversarial training generates adversarial examples during training and mixes them into the training data with correct labels, so the model learns to get them right.
  • Goodfellow, Shlens and Szegedy introduced it in 2014 as a regularizer using the fast gradient sign method; Madry and colleagues reformulated it in 2017 as a min-max problem solved with projected gradient descent.
  • NIST names it one of three evasion defences that have proven resilient, alongside randomized smoothing and formal verification, and resistance grows as the attacks used in training get stronger.
  • The cost is real. NIST says it usually reduces accuracy on clean data and is expensive, and Madry's team found a network needs more capacity to be robust than to classify cleanly.
  • NIST calls a model that resists evasion without losing accuracy an open problem, not something adversarial training has solved.

Adversarial training is a defence that generates adversarial examples during training and mixes them into the training data with their correct labels, so the model has to learn to get them right. Goodfellow, Shlens and Szegedy introduced it in 2014 as a regularizer; Madry and colleagues reframed it in 2017 as a robust-optimization problem. NIST lists it as one of three evasion defences that have “proven to be resilient” against a field where most proposed defences have not.

How it works

Goodfellow and colleagues trained on a weighted mix of the ordinary loss and the loss on a fast gradient sign method adversarial example built from the current model. Because that example is regenerated from whatever the model currently is, training keeps chasing a moving target: as the model changes, so does the example it has to get right.

Madry and colleagues reframed this as a single min-max problem: minimise, over the model’s parameters, the worst loss an attacker could produce by perturbing each input within some allowed limit. Solving that inner “worst loss” part exactly is intractable, so they use projected gradient descent as what they call a reliable, “ultimate” first-order adversary — it does not find the true worst case, but training against it produces broad robustness. NIST’s framing follows this lineage and credits Madry’s group with developing it further.

Why it matters

Most proposed defences against evasion have not survived a determined attacker; NIST recounts entire classes of detection and gradient-masking defences being broken soon after publication. Adversarial training, randomized smoothing and formal verification are the three NIST says have held up. Adversarial training’s mechanism is also its argument: resilience grows as the attacks used in training get stronger, so the defence keeps a path to improve rather than depending on one attack staying unbroken.

In practice

Even the strongest published results fall well short of immunity. Madry and colleagues’ MNIST model, trained against a PGD adversary, held above 89% accuracy against their strongest test-suite adversaries; their CIFAR10 model, a harder task, held only 46% against the same class of adversary. NIST’s account of why so few defences reach even that bar names specific failures: Carlini and Wagner defeated ten proposed detection methods, the Obfuscated Gradients attack broke defences relying on masked gradients, and Tramèr and colleagues circumvented thirteen more by designing attacks adapted to each one.

Trade-offs

The gain is not free. NIST reports that adversarial training “usually comes at the cost of decreased model accuracy on clean data” and is “expensive due to the iterative generation of adversarial examples during training.” Madry and colleagues give a structural reason for the accuracy cost: a network needs more capacity to draw a robust decision boundary than it needs to draw one that only has to separate ordinary, unperturbed data, so part of a fixed-size model’s capacity goes to robustness instead of clean-data accuracy.

The cost does not scale evenly either. NIST notes that adversarial training, widely used for a single input type such as images, “may become prohibitively expensive as the number of modality combinations increases.” Training a model that reads several modalities together on clean data alone does not make it more robust against an attack on just one of them, either.

None of this makes adversarial training a stopgap. NIST names it alongside randomized smoothing and formal verification as the field’s most resilient options, while also stating plainly that designing a model that resists evasion while keeping its accuracy “remains an open problem.”

Where definitions disagree

NIST and MITRE ATLAS classify adversarial training differently. NIST treats it as one of three named, first-class defence categories against evasion, alongside randomized smoothing and formal verification. ATLAS does not give it a mitigation entry of its own: it appears as one listed technique inside AML.M0003, “Predictive AI Model Hardening”, a mitigation that also lists robust architectures, defensive distillation and certified robustness methods. The underlying technique is the same; whether it counts as its own category or as one instance of a broader hardening mitigation depends on which framework is doing the counting.

Questions and answers

Does adversarial training make a model immune to adversarial examples?

No. NIST calls designing an ML model that resists evasion while maintaining accuracy an open problem, not a solved one. Adversarial training's own record shows the limit: Madry and colleagues' CIFAR10 model, one of the strongest published results in 2017, held only 46% accuracy against the adversary it was built to resist. Training on today's attacks raises the bar for tomorrow's, but it does not remove the possibility of a stronger attack finding a new gap.

Why does adversarial training hurt accuracy on clean data?

NIST reports this cost as usual, not incidental, across the class of mitigations that includes adversarial training. Goodfellow and colleagues' own explanation is capacity: they judged adversarial training "only clearly useful when the model has the capacity to learn to resist adversarial examples", and Madry and colleagues later confirmed experimentally that a network needs more capacity to be robust than to classify cleanly. A fixed-size model spends some of that capacity on robustness instead of on the clean-data task.

Is adversarial training the same on every kind of input?

No. NIST notes that adversarial training, common in single-modality systems, can become prohibitively expensive as the number of modality combinations grows, and that training a multimodal model on clean data alone does not make it more robust against an attack on a single modality. The technique that works well for one image classifier does not simply scale to a model that reads images, audio and text together.

Sources

  1. Explaining and Harnessing Adversarial ExamplesarXiv, 20 Dec 2014
  2. Towards Deep Learning Models Resistant to Adversarial AttacksarXiv, 19 Jun 2017
  3. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  4. MITRE ATLAS, AML.M0003 Predictive AI Model Hardening (collection 2026.08)MITRE

Guides that use this term