What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Adversarial machine learning

Adversarial machine learning is the field that studies attacks against machine learning systems which exploit the statistical, data-based nature of those systems, such as poisoned training data, manipulated inputs and extraction of private information. It also studies how to design models that withstand such attacks, a property often known as robustness.

Last reviewed

Key points

  • Adversarial machine learning studies attacks that exploit the statistical, data-based nature of machine learning, and how to design models that withstand them.
  • The attacks span the whole life cycle, from poisoning training data to fooling a deployed model or querying it for secrets.
  • It needs an adversary. NIST puts failures "outside the context of adversarial use, such as inaccuracy" outside the field.
  • It is not the whole of AI security. Attacks on the platform a model runs on belong to ordinary cybersecurity, and securing an AI system takes both.
  • NIST says the field is "neither a complete solution to, nor a subset of" AI safety or the other trustworthy-AI characteristics.

Adversarial machine learning studies the attacker, not everything that can go wrong with a model.

What the field studies

NIST’s 2025 report says the field studies attacks that exploit the “statistical, data-based nature of ML systems.” The field is also concerned with designing systems that withstand them, “a property often known as robustness.”

The attacks reach every stage of a model’s life. NIST lists “adversarial manipulation of training data,” as in data poisoning; “adversarial inputs” that degrade performance, such as adversarial examples; and interactions that pull sensitive information out of the model’s training data, as in membership inference.

NIST sorts attacks by what they break: availability, integrity or privacy. For generative AI it adds a fourth goal, misuse enablement: getting round the restrictions an owner puts on a model, as in jailbreaking.

NIST dates early evasion attacks to 1988, and to 2004 against spam filters.

Why it matters

Few defences come with guarantees. NIST says there are “no information-theoretic security proofs for the widely used ML algorithms in modern AI systems,” meaning no mathematical proof that they are secure, so many mitigations are adopted “because they appear to work in practice.”

NIST also records the field’s own failures: “a history of publishing defenses evaluated under relatively weak adversarial models that are subsequently broken by more powerful attacks.”

Where the field ends

The boundary is the adversary. NIST puts design and implementation flaws that cause failures “outside the context of adversarial use, such as inaccuracy” outside the field. A model that is simply wrong is not an adversarial machine learning problem.

It is not the whole of AI security. NIST leaves attacks on the platforms a model runs on to “traditional cybersecurity taxonomies,” and says securing AI systems needs both.

Nor is it a branch of AI safety. NIST says robustness to these attacks “may play an important role” in safety, but the field is “neither a complete solution to, nor a subset of” AI safety or the other characteristics of trustworthy AI.

Questions and answers

What is adversarial machine learning?

Adversarial machine learning is the field that studies attacks against machine learning systems which exploit how those systems learn from data, such as poisoned training data, manipulated inputs and extraction of private information, and how to build models that withstand them. NIST's 2025 report on the field says that property is "often known as robustness".

Is adversarial machine learning the same as AI security?

No. Adversarial machine learning covers attacks that exploit the statistical nature of a model. NIST places attacks against the platforms AI systems run on within traditional cybersecurity, and says managing the security of AI systems "will require combining mitigations from the field of AML with best practices for the development of secure software from the field of cybersecurity."

Does adversarial machine learning cover AI mistakes and hallucinations?

Not when there is no attacker. NIST excludes design and implementation flaws that cause failures "outside the context of adversarial use, such as inaccuracy" from the scope of the field. An attacker who deliberately induces a wrong output is in scope.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025

Guides that use this term