What matters in AI.

Subscribe

Learn / AI security

Guide · AI security

White-box vs black-box attacks

White-box and black-box attacks on AI models differ in what the attacker knows. A white-box attacker has full knowledge of the model, including its architecture and parameters, and can compute gradients to find an attack. A black-box attacker has little or no knowledge and can only query the model and observe its outputs. Gray-box attacks sit between the two.

Last reviewed

The difference is what the attacker knows. NIST’s 2025 taxonomy of adversarial machine learning says a white-box attacker operates “with full knowledge about the ML system”. A black-box attacker has “minimal, and sometimes no knowledge at all” and may be able to query the model but has “no other information about how the model is trained”. NIST uses this as one of the dimensions for sorting attacks, separate from what the attack does. An adversarial example, a privacy attack or a jailbreak can each be attempted white-box or black-box.

Side by side

White-box Black-box
What the attacker has The model’s architecture and parameters, often more An interface: send an input, read the output
How the attack is found Gradients: the search is an optimisation problem Queries, a substitute model, or both
Cost Gradient steps on the attacker’s own copy: under 100 or 1000 in Carlini and colleagues’ 2019 image example Queries to the target: often orders of magnitude more in the same example
Who uses it Defenders testing the worst case; anyone with the weights Attackers facing a deployed service

What white-box access buys

With the weights, finding an attack is ordinary optimisation. In NIST’s words, the adversary can “compute gradients relative to the model’s loss function”. The search looks for a change to the input that flips the output, “often with constraints on properties such as the perceptibility or size of the perturbation”. NIST lists fast gradient sign method and projected gradient descent among white-box evasion attacks.

The cost is computation, not access. Carlini and colleagues wrote in 2019 that white-box attacks on the CIFAR-10 and ImageNet image benchmarks “generally converge in under 100 or 1000 iterations”. That figure is for a perturbation that changes each pixel by at most 8/255 of its range and a step size of 1. Black-box attacks “often take orders of magnitude more queries”.

NIST is candid that white-box attacks “operate under very strong assumptions”. It gives the reason for studying them anyway: “to test the vulnerability of a system against worst-case adversaries and to evaluate potential mitigations.” NIST’s white-box definition also covers the adaptive attack, in which “knowledge of the mitigations applied to the model or the system is explicitly tracked”.

What a black-box attacker can still do

A black-box attacker works through the interface. For black-box model evasion, NIST splits the attacks by how much the interface reveals. In score-based attacks the attacker sees the model’s confidence scores. A popular method, NIST says, uses them to estimate the model’s gradients without computing derivatives. In decision-based attacks the attacker can “only obtain the final predicted labels”. NIST says the main challenge is reducing the number of queries, and that recent techniques can evade classifiers with “a relatively small number of queries, typically less than 1000”.

Brendel, Rauber and Bethge argued in 2018 that decision-based attacks matter most in practice. They need only what a real service shows, and they are “more robust to simple defences than gradient- or score-based attacks”. So hiding confidence scores removes one class of attack, not all of them.

The other route is to turn a black-box problem into a white-box one. Papernot and colleagues trained a local substitute on inputs labelled by the target, attacked the substitute white-box, and sent the results to the real service. A handwritten-digit classifier the authors trained through MetaMind’s service misclassified 84.24% of them. Digit classifiers they trained on Amazon’s and Google’s services misclassified 96.19% and 88.94%, attacked through logistic regression substitutes built from 800 queries. All three results are at a perturbation of ε = 0.3. This works because of transferability. NIST adds that model extraction is “often not an end goal but a step toward other attacks”. As the weights and architecture become known, attackers can launch attacks “typical for the white-box or gray-box settings”.

Gray-box: the range in between

The two labels are ends of a range. NIST describes gray-box attacks in between: an attacker who knows the architecture but not the parameters, or the model but not the training data, or who has data “distributed identically to the training data”. Carlini and colleagues likewise describe “varying degrees of black-box access”: a query budget, probabilities or only the predicted class.

Knowing what your system reveals tells you where on that range an attacker starts. Returning scores, publishing the architecture, or building on a public base model all move the attacker toward white-box.

Why defences are tested white-box

A defence that only holds against black-box attackers is relying on secrecy. Carlini and colleagues’ guide to evaluating robustness says a threat model should “assume that the adversary has complete knowledge of the inner workings of the defense”. It is “not reasonable” to assume the defence can be kept secret, “even in black-box threat models”. That is Kerckhoffs’ principle from cryptography. White-box robustness “implies robustness to black-box adversaries, but not the other way around”.

The same guide gives a test. White-box attacks are “a strict super-set of black-box attacks” and should do strictly better. If gradient-free attacks do better, that “often indicates that the defense is somehow masking gradient information”, a pattern called gradient masking. Athalye, Carlini and Wagner found this in 7 of 9 non-certified defences at ICLR 2018 that claimed white-box robustness. They fully circumvented 6 and partially circumvented a seventh.

How it plays out for language models

For generative AI, NIST’s black-box end includes “systems with deliberately obscured or misleading information”. The white-box end is easier to reach than it looks, because open-weight models exist. Zou and colleagues built an adversarial suffix by gradient search on two open Vicuna models. The suffix then induced objectionable content in the public interfaces to ChatGPT, Bard and Claude. Transfer worked much better on the GPT-based models, which the authors say may be because Vicuna was trained on ChatGPT outputs. NIST draws the lesson: because such universal triggers transfer, open-weight models, “for which there is ready white-box access”, are “feasible attack vectors” against closed systems “in which only API access is available”.

Black-box attacks on language models need not be expensive either. PAIR, from Chao and colleagues, uses one language model to attack another with “only black-box access” and “often requires fewer than twenty queries”.

Where definitions disagree

What counts as “full knowledge” varies. NIST’s general definition of white-box includes “the training data, model architecture, and model hyperparameters”. Carlini and colleagues define it as “complete knowledge of the model and its parameters” and do not mention training data. The difference matters for open-weight models. By Carlini’s wording, an attacker with the weights but not the training data is white-box. NIST’s gray-box description covers exactly that attacker: one who knows “the model and its parameters but not the training data”.

The number of categories varies too. NIST names three, with gray-box between the ends. Wiz, a security vendor, names two, white box and black box. It also states that “Most real-world adversarial AI incidents are black box attacks” but does not give a source. NIST makes a narrower claim: black-box attacks are “the most practical”, because they use “system interfaces readily available for normal use”.

Questions and answers

Are black-box attacks weaker than white-box attacks?

In principle, yes: Carlini and colleagues call white-box attacks "a strict super-set of black-box attacks". In practice the gap can be small. Their 2019 guide says the decision-based Boundary Attack "often rivals the best white-box attacks" in how small a change it needs, at the cost of many more queries. A black-box attacker can also build a white-box attack on a substitute model and transfer it.

Is a model safe from adversarial attacks if its weights are private?

No. Private weights make a direct white-box attack harder, but black-box and transfer attacks still work. Papernot and colleagues attacked classifiers hosted on MetaMind, Amazon and Google services using only the predicted labels, by training a substitute model and transferring attacks from it. For language models, NIST notes that universal jailbreak triggers transfer between models, which makes open-weight models "feasible attack vectors" against closed systems that offer only API access.

Sources

  1. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  2. On Evaluating Adversarial RobustnessarXiv (Carlini et al.), 20 Feb 2019
  3. Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial ExamplesarXiv (Athalye et al.), 31 Jul 2018
  4. Practical Black-Box Attacks against Machine LearningarXiv (Papernot et al.), 19 Mar 2017
  5. Decision-Based Adversarial Attacks: Reliable Attacks Against Black-Box Machine Learning ModelsarXiv (Brendel et al.), 16 Feb 2018
  6. Universal and Transferable Adversarial Attacks on Aligned Language ModelsarXiv (Zou et al.), 20 Dec 2023
  7. Jailbreaking Black Box Large Language Models in Twenty QueriesarXiv (Chao et al.), 13 Oct 2023
  8. Adversarial AI: Understanding and Mitigating Machine Learning ThreatsWiz

Terms in this guide