What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Adversarial example

An adversarial example is an input modified so that a machine learning model misclassifies it or misbehaves at deployment time. The modification targets how the model draws its decision boundary rather than any flaw in the surrounding code, and in the best-known cases it is small enough that a person still reads the input correctly.

Last reviewed

Key points

  • An adversarial example is an input altered so a model gets it wrong at deployment time, while the surrounding code is correct and the model is the one that passed its tests.
  • In the founding research the alteration is invisible to a person, but imperceptibility is a convention of the literature rather than part of the definition, and NIST's glossary drops it.
  • An attacker does not need access to the target. Adversarial examples transfer, so one generated against a model the attacker controls can work on the model they cannot reach.
  • Detection is not a way out. NIST holds that detecting adversarial examples is equivalent to robust classification, and ten published detectors were defeated by constructing new loss functions.
  • Adversarial training is the defence with the longest record, and it costs accuracy on clean data. NIST calls resisting evasion while keeping accuracy an open problem.

An adversarial example is an input someone changed so that a model gets it wrong. NIST’s glossary keeps to what can be observed: “a modified testing sample that induces misclassification or misbehavior of a machine learning model at deployment time.” In the 2013 work that named it, the change is invisible — Szegedy and colleagues made a network misclassify an image “by applying a certain imperceptible perturbation”. Nothing around the model is broken. The code is correct, the model is the one that was tested, and the answer is wrong on demand.

How an adversarial example is made

An adversarial example is made by searching for a change to the input that moves the model’s output where the attacker wants it, within a limit on how large that change may be. Method follows access. Given the architecture and parameters, gradients make the search an ordinary optimisation problem — what both 2013 founding papers required.

With query access only, NIST lists zeroth-order optimisation, discrete optimisation, Bayesian optimisation and transferability, the last of which removes the need for access to the target: generate the example white-box against a model you control, then send it to the one you do not.

Constrained inputs are the defender’s one piece of luck. An image perturbation can be any small vector; one applied to a PDF or a network packet must leave a valid PDF or packet. NIST reports that attacks built for continuous domains “result in infeasible adversarial examples” there.

Where definitions disagree

The artefact and the attack are separate things, and the law merges them. NIST is precise — evasion is the attack, the adversarial example is the input it produces — while Article 15(5) of the EU AI Act names one item, “adversarial examples or model evasion”.

Imperceptibility is a convention, not a requirement. The founding papers built the term on perturbations a person cannot see, and NIST’s body text says computer-vision examples are “often designed to be imperceptible to humans”, but its glossary drops the condition. Against a commercial phishing detector, real attackers used “relatively simple methods for evasion, such as image cropping, masking, or blurring techniques.”

“Adversarial machine learning” is wider than this topic, covering poisoning and privacy attacks as well. The nearest neighbour is model poisoning, which corrupts the model before deployment; an adversarial example leaves a correct model intact.

Questions and answers

What is the difference between an adversarial example and a prompt injection?

An adversarial example is an input perturbed so that a model's prediction changes, and the target is usually a classifier making a decision. A prompt injection is text that a language model reads as an instruction when it was supposed to be data. NIST's 2025 taxonomy keeps them apart structurally: evasion attacks and adversarial examples sit in its predictive AI taxonomy, while direct prompting attacks and indirect prompt injection are separate classes in its generative AI taxonomy. The boundary is not perfectly clean even there, because NIST discusses an ASCII-art attack on chatbots inside its section on text adversarial examples.

Are adversarial examples always invisible to humans?

No. Imperceptibility is a property of the most-cited research, not a requirement of the term. NIST says adversarial examples are "often designed to be imperceptible to humans" in computer vision, and its glossary asks only that a modified testing sample induce misclassification or misbehaviour at deployment time. The attacks NIST records from the field are mostly visible: image cropping, masking and blurring against a phishing detector, and masks, wigs and deepfakes against a face verification service.

Can you detect adversarial examples and reject them?

Not reliably, and there is a proof of why. Tramèr showed a reduction from detection to classification: a detector that is robust at some perturbation distance yields a classifier that is robust at half that distance. NIST draws the conclusion that designing techniques to detect adversarial examples "is equivalent to robust classification, which is inherently difficult to solve". A detector that reliably spots manipulated inputs is a model that reliably resists them, so the easier-looking problem is the same problem.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. Intriguing properties of neural networksarXiv, 21 Dec 2013
  3. Evasion Attacks against Machine Learning at Test TimearXiv, 21 Aug 2017
  4. Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection MethodsarXiv, 20 May 2017
  5. Detecting Adversarial Examples Is (Nearly) As Hard As Classifying ThemarXiv, 24 Jul 2021
  6. Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, Robustness and CybersecurityEuropean Union, 12 Jul 2024

Guides that use this term