What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Transferability

Transferability is the property that an adversarial example crafted against one machine learning model often also fools a different model trained for the same task, even one with a different architecture or different training data. Transferability lets an attacker build a malicious input on a model they control and use it against a model they cannot inspect.

Last reviewed

Key points

  • In the research that defines it, transferability is a property of adversarial examples. The attack that exploits it is a transfer attack.
  • It means an attacker does not need the target's internals. They build the example on a model they control, then send it to the one they cannot inspect.
  • Transfer is far from guaranteed. In one 2013 table of small digit classifiers trained on the same data, examples made for one fooled the others between 2% and 87% of the time; random noise alone reached 23%.
  • With the methods of 2016, making an image model wrong transferred far more easily than making it give a chosen wrong answer.
  • It reaches language models. In 2023 a jailbreak suffix built on open models raised GPT-3.5's rate of harmful answers from 1.8% to 47.4%; Claude 2 went from 0% to 1.8%.

Transferability lets an attacker fool a model whose internals they have never seen by attacking a different model instead. An adversarial example built for one model often fools another, across architectures and training sets, “so long as both models were trained to perform the same task”, in the words of Papernot, McDaniel and Goodfellow.

How transferability works

A transfer attack typically has three steps, NIST says. Get a surrogate model that does the target’s job. Craft adversarial examples against the surrogate. Send them to the target.

NIST gives two examples of getting the surrogate. Papernot and colleagues trained theirs on answers queried from the target, which is also how model extraction builds a copy. Other papers trained several models of their own, an ensemble, without explicitly querying the target.

Why examples carry over is still being worked out. Goodfellow, Shlens and Szegedy argued that adversarial examples “occur in broad subspaces”, not fine pockets, so one misclassified by one classifier has “a fairly high prior probability” of being misclassified by another. Tramèr and colleagues found two digit-recognition networks shared a significant fraction of those subspaces.

Why it matters

Transferability means the attacker does not need the target’s internals, only a way to send it inputs. Of the methods NIST lists for black-box settings, where the attacker can only query the target, transferability is the one that does its search on a different model.

Hosted services are not exempt. Papernot and colleagues trained digit classifiers through MetaMind, Amazon and Google, then attacked them from surrogates without knowing the architectures. MetaMind’s model misclassified 84.24% of the examples; with a logistic-regression surrogate and 800 queries, Amazon’s and Google’s misclassified 96.19% and 88.94%.

NIST says the same holds for language models: open-weight models are a feasible route to attack closed ones served only through an API.

How often it works

Transfer is partial, and the rate depends on what the attacker needs and how big a change they may make.

The Papernot figures used a large change: each pixel could move by up to 0.3 on a scale of 0 to 1. The targets were digit classifiers the authors had trained through each service, not the services’ own products.

In their 2013 paper, Szegedy and colleagues fed examples made for one small model trained on MNIST, a dataset of handwritten digits, to five others trained on the same data. The others got them wrong between 2% and 87.1% of the time, depending on the pair. Random noise, mostly larger than the adversarial changes, caused between 0% and 22.7%.

Making a model wrong transferred far better than making it wrong in a chosen way. On ImageNet, a large photo dataset, Liu and colleagues found examples that only had to cause a mistake transferred readily. Examples aimed at a chosen label “almost never transfer with their target labels” with the methods of 2016: 0% to 4% between single models. Building each example against four models at once raised the rate on a fifth, unseen model to between 11% and 46%, at the cost of larger changes to the image.

Language models show the same unevenness. Zou and colleagues optimised a jailbreak suffix on open Vicuna and Guanaco models and tested it on 388 harmful requests in 2023. One suffix raised GPT-3.5’s rate of harmful answers from 1.8% to 47.4%. Claude 2 went from 0% to 1.8%.

Transfer is not a law either. Tramèr and colleagues built an artificial digit task on which a linear and a quadratic model, each easy to fool directly, did not share adversarial examples. They left open whether any real dataset behaves that way, and concluded that defences against transfer attacks may be possible even for models vulnerable to direct attack.

Where definitions disagree

Papers define transferability as a property. Papernot, McDaniel and Goodfellow call it “the property that some adversarial samples produced to mislead a specific model” can mislead others. NIST uses the word both ways. It lists transferability among black-box methods and opens its section on the subject with “Another method for generating adversarial attacks”, while calling attack transferability “an intriguing phenomenon” in the same section.

NIST carries the method sense beyond model evasion. When the attacker knows only part of the target, it calls transferability the most popular method for one kind of data poisoning: building poisoned training samples against a surrogate and using them on the target.

This page follows the property sense: transferability is the property, and a transfer attack is what exploits it.

Questions and answers

Is transferability an attack?

Not in the sense this page uses. Papernot, McDaniel and Goodfellow define transferability as a property: some adversarial examples made to mislead one model also mislead others. A transfer attack exploits it by building the example on a surrogate model and sending it to the target. NIST uses the word for both, calling transferability a method and attack transferability a phenomenon.

Does keeping my model private stop adversarial examples?

Not reliably. Transferability means an attacker can craft adversarial examples on a model they own and send them to yours. Papernot and colleagues trained digit classifiers through MetaMind, Amazon and Google, built surrogates from the labels those services returned, and most examples crafted on the surrogates fooled the hosted models. Keeping a model private raises the attacker's cost; it does not remove the attack.

Do jailbreaks transfer between language models?

Some do, unevenly. In 2023 Zou and colleagues optimised a suffix on open models and it raised GPT-3.5's rate of harmful completions from 1.8% to 47.4% of 388 harmful requests, while Claude 2 went from 0% to 1.8%. NIST notes that this makes open-weight models a feasible route for attacking closed models available only through an API.

Sources

  1. Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial SamplesarXiv, 24 May 2016
  2. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  3. Intriguing properties of neural networksarXiv, 21 Dec 2013
  4. Explaining and Harnessing Adversarial ExamplesarXiv (ICLR 2015), 20 Dec 2014
  5. Practical Black-Box Attacks against Machine LearningarXiv (ASIA CCS 2017), 8 Feb 2016
  6. Delving into Transferable Adversarial Examples and Black-box AttacksarXiv (ICLR 2017), 8 Nov 2016
  7. The Space of Transferable Adversarial ExamplesarXiv, 11 Apr 2017
  8. Universal and Transferable Adversarial Attacks on Aligned Language ModelsarXiv, 27 Jul 2023

Guides that use this term