What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Model evasion

Model evasion is an attack technique, catalogued by MITRE ATLAS as AML.T0015, in which an adversary crafts or presents input so a deployed AI model misreads it at inference time. The input may be a perturbed image the model should classify correctly, or a synthetic one, such as a deepfake, presented as authentic. ATLAS marks it Realized.

Last reviewed

Key points

  • Model evasion (ATLAS AML.T0015) covers two routes — crafting adversarial data a model misreads, or generating a deepfake and presenting it as genuine. ATLAS marks it Realized, meaning documented incidents exist, not only research.
  • The same technique appears under three ATLAS tactics depending on the goal — Initial Access when a deepfake defeats a biometric check, Defense Evasion when a file slips past a malware detector, Impact when a deployed model is simply made to fail.
  • Evasion targets a model at deployment time; data poisoning corrupts what it learned during training. NIST keeps the two apart by when the attack happens, not by what it costs.
  • Fielded evasion is cruder than the research literature. NIST found attackers beating a commercial phishing detector with image cropping, masking and blurring — not gradient-based optimization.

How it works

ATLAS’s definition names two paths to the same result. An attacker can craft adversarial data a model reads wrong — a perturbed image, a mutated domain name, a modified malware sample — or generate a deepfake and present it as genuine input instead. Both count as evasion, because both make a deployed model misperform at test time.

Where the technique sits in an attack depends on the goal, not the method. ATLAS case studies that employ it span three tactics: Initial Access, when a faked identity gets past a biometric check; Defense Evasion, when a crafted file slips past a malware or spam detector already watching for it; and Impact, when the point is simply making a deployed model fail at its job.

NIST separates evasion from data poisoning by timing, not technique: evasion targets a model at deployment time, after training is finished, while poisoning corrupts what the model learns while it is still training.

Why it matters

Fielded evasion looks nothing like the literature’s optimized perturbations. NIST reports that when researchers manually reviewed the adversarial examples a commercial phishing detector had flagged as uncertain, the attackers had not used gradient-based optimization at all — they cropped, masked or blurred a brand logo by hand. The cheap method already worked, so nobody needed the expensive one.

That case also shows why evading one check does not mean evading the system. The phishing pages that beat the detector’s brand-similarity model were still caught by its other components, because the system did not rely on one signal alone. NIST is also careful to label several of its catalogued cases — the Palo Alto and Cylance exercises among them — as researcher demonstrations that never became attacks in the wild, which is a different claim from the tax-fraud case, where the fraud was real.

In practice

Three ATLAS-catalogued cases show the range, from a researcher’s proof of concept to real fraud.

Palo Alto Networks’s own AI research team bypassed a CNN-based Domain Generation Algorithm detector using a generic domain-name mutation technique effective against most ML-based DGA detectors — an internal exercise, not an attack anyone suffered.

Researchers evading ProofPoint’s email protection system went a step further: they built a copy-cat model, learned its scoring from the outside, and used that to craft emails the real ProofPoint system scored as clean. That is transferability applied to spam filtering rather than image classification — the target model never had to be touched directly.

The tax-fraud case is not an exercise. Two individuals in China used a camera hijack attack to defeat the live facial recognition system on the Shanghai government’s tax portal, impersonating a victim to file fraudulent invoices through a fake shell company from 2018 onward, and fraudulently collected 77 million US dollars before being caught.

And the phishing-detector case argues for defense in depth. The detector was an ensemble: one component scored brand-logo similarity, others checked different signals. Attackers beat the logo-similarity component with hand-modified logos, but the phishing pages were still flagged by the system’s other components. Evading one check is not the same as evading the system built around it.

Where definitions disagree

ATLAS and NIST draw the boundary of “evasion” in different places. NIST’s evasion attack is specifically about adversarial examples — samples perturbed so a model’s classification changes. ATLAS’s Evade AI Model names a second, separate path: generating a deepfake and presenting it as genuine input, with no perturbation of a real sample involved. A silicone mask worn to defeat a face scanner is not an adversarial example by NIST’s definition — nothing about the scanner’s decision boundary was probed or exploited — but ATLAS still counts it as evasion, because the result is the same: a deployed model accepts input it should have rejected. This page follows ATLAS’s broader scope, since deepfake attacks on biometric systems are exactly the sort of evasion the real-world case studies record.

Questions and answers

What is the difference between model evasion and an adversarial example?

Model evasion is the ATLAS technique — the adversary's action of getting a deployed model to fail at inference time. An adversarial example is one way to do it: an input perturbed so the model misclassifies it. ATLAS's definition of Evade AI Model names crafting adversarial data as one path and generating a deepfake as the other, so evasion is the broader category and an adversarial example is one of its artifacts, not a synonym for it.

Does model evasion require access to a model's training data?

No. NIST separates evasion from data poisoning by when the attack happens: evasion targets a model already deployed, at inference time, while poisoning requires the attacker's data to get into the training set before the model is trained. An attacker who only ever sees a model's live predictions, never its training data, can still evade it — the ProofPoint case did exactly that, using a copy-cat model trained on the attacker's own data to learn what the real system would accept.

Why do real-world evasion attacks look cruder than the research literature's optimized attacks?

Because the cruder method already works. NIST found that when a commercial phishing detector flagged 4,600 samples as uncertain, the 100 that were adversarial examples had been made with image cropping, masking and blurring, not gradient-based optimization. An attacker who only needs to clear one specific deployed check has no reason to compute a gradient if cropping a logo does the job.

Sources

  1. MITRE ATLAS, technique AML.T0015 Evade AI Model (collection 2026.09)MITRE
  2. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  3. Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, Robustness and CybersecurityEuropean Union, 12 Jul 2024

Guides that use this term