What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Adaptive attack

An adaptive attack is an attack designed against one specific defence, by an attacker who knows how that defence works and changes their method to defeat it. It is the standard test of defences against adversarial examples, and researchers argue it should be for jailbreak and prompt injection defences too.

Last reviewed

Key points

  • An adaptive attack is built against the specific defence in place, by an attacker who knows how the defence works.
  • Carlini and colleagues call an evaluation against non-adaptive attacks "of very limited utility"; running existing attacks with default settings is not enough.
  • Tramèr and colleagues circumvented all 13 adversarial-example defences they studied, even though nearly all had already run their own adaptive evaluation. Adaptive attacks "cannot be automated".
  • Nasr and colleagues bypassed 12 recent defences against jailbreaks and prompt injection, most with attack success above 90%, though most had reported near-zero success against static or weak attacks.
  • Passing an adaptive attack does not prove a defence works. It only fails to show it is broken.

How it works

Carlini and colleagues define adaptive adversaries as ones “adapted to the specific details of the defense” that try to invalidate its robustness claims. The term comes from general security, for an attacker who “adapts to what the defender has done”.

Tramèr and colleagues describe how they attacked adversarial example defences. In most cases:

  1. Read the paper and the source code, looking for reasons the defence might still be beaten.
  2. Write a score for the attack to push on (a loss function) that measures progress against this defence and is easy to optimise.
  3. Pick a way to optimise that score, run the attack, learn from what fails, and repeat.

Their advice is to keep it simple: attack the full defence end to end, and where a defence has many parts, look for the one or two that matter. Make sure a better score always means a stronger attack.

A common reason a defence needs this is gradient masking. The defence breaks the gradient that standard attacks follow. An adaptive attacker who knows this can recover the gradient another way.

Why it matters

Carlini and colleagues write that an evaluation against non-adaptive attacks “is of very limited utility”, and that with existing attacks at default settings “there is a good chance these attacks will work sub-optimally or even fail against a new defense”.

Tramèr and colleagues studied 13 defences from ICLR, ICML and NeurIPS. Nearly all had run their own adaptive evaluation, often by reusing an attack that had broken an earlier defence. All 13 were circumvented. The tools for a stronger attack already existed. The problem, they conclude, was “methodological rather than technical”: evaluators did not tailor them to the defence.

The same pattern has reached language models. Nasr and colleagues tested 12 recent defences against jailbreaking and prompt injection. Most had reported near-zero attack success rate, tested against static sets of attack prompts or weak optimisation attacks. Adaptive attacks pushed success above 90% for most of them.

In practice

NIST’s guidance on adversarial machine learning says mitigations against adversarial examples “need to be evaluated against strong adaptive attacks”. Its definition of a white-box attack covers adaptive attacks, “in which knowledge of the mitigations applied to the model or the system is explicitly tracked”.

Whether adaptive attacks can be automated is open. Tramèr and colleagues found no single strategy worked on every defence, and state that adaptive attacks “cannot be automated”. Against LLM defences, Nasr and colleagues found automated search and reinforcement learning attacks can systematically bypass defences, though they call them less reliable than attacks like PGD on images. On a selected subset, their search attack succeeded 69% of the time and human red-teamers collectively succeeded every time. They encourage industry labs to keep up human AI red teaming, with the red-teamers given details of any defences in the system.

Tramèr and colleagues also warn against copying their attacks as templates. For any fixed attack, “it is possible to build a non-robust defense that prevents that attack”.

Where definitions disagree

The phrase is used in two different senses.

Adapted to the defence. Carlini, Tramèr, Nasr and NIST mean an attack designed for one specific defence, by an attacker who knows how it works.

Adapted to the input. Wei, Haghtalab and Steinhardt’s jailbreak study used “adaptive attack” for an adversary who “can choose the attack based on the prompt”. It counted a success “if any of the 28 evaluated attacks succeed”. That is picking the best of a fixed menu for each prompt, not designing a new attack against a defence. It still made the point: on the authors’ curated set of 32 harmful prompts, Claude v1.3 resisted every roleplay attack, yet was “100% vulnerable to an adaptive attack”.

When a paper reports results against an adaptive attack, check which one it means.

Questions and answers

What is the difference between an adaptive attack and a white-box attack?

A white-box attack describes what the attacker knows: the model, its weights, its training. An adaptive attack describes what the attacker does with knowledge of the defence: change the attack to beat it. NIST says its white-box definition "encompasses the notion of adaptive attacks". An adaptive attack can still use only queries, as when Nasr and colleagues' human red-teamers beat defences with limited information and feedback.

Can adaptive attacks be automated?

Not fully. Tramèr and colleagues found no single strategy worked on all 13 defences they circumvented and conclude that adaptive attacks "cannot be automated". Nasr and colleagues found automated search and reinforcement learning attacks promising against LLM defences, but on a selected subset the human red-teamers, taken together, still did better.

Sources

  1. On Evaluating Adversarial RobustnessarXiv (Carlini et al.), 20 Feb 2019
  2. On Adaptive Attacks to Adversarial Example DefensesarXiv (Tramèr et al., NeurIPS 2020), 23 Oct 2020
  3. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
  4. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsarXiv (Nasr et al.), 10 Oct 2025
  5. Jailbroken: How Does LLM Safety Training Fail?arXiv (NeurIPS 2023), 5 Jul 2023

Guides that use this term