What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Attack success rate

Attack success rate (ASR) is the fraction of test cases in which an attack achieved the attacker's goal against an AI model or agent. An attack success rate describes a whole test, not the system alone: it depends on which attacks were tried, how many attempts each got, which tasks were tested and how success was judged.

Last reviewed

Key points

  • Attack success rate (ASR) is the share of test cases in which an attack achieved its goal. Benchmarks such as HarmBench and AgentDojo report it.
  • It measures the attacks as much as the system. US and UK government red teamers raised the rate against one AI agent from 11 percent to 81 percent on held-back tasks, by changing only the attacks.
  • Retries change it. Allowing 25 tries per attack raised NIST's average across five attacker goals from 57 percent to 80 percent.
  • A low rate does not mean a real attacker will usually fail. Attacks built for each defence bypassed all 12 recently proposed defences in one study, most at above 90 percent, though most had reported near-zero rates.

What an attack success rate depends on

Run a set of test cases, count those where the attacker got what they wanted, and divide by the total. Four choices sit behind that fraction.

The attacks. NIST’s Center for AI Standards and Innovation (CAISI) tested an agent built on a model AgentDojo had found among the most resistant to agent hijacking. With red teamers from the UK AI Security Institute, CAISI built new attacks on some tasks in one environment and scored them on tasks it had held back. The rate went from 11 percent for the strongest existing attack to 81 percent for the strongest new one. The agent was the same.

The attempts. Many evaluations give each attack one try, but a model’s output varies between runs. CAISI tried the attack for each of five attacker goals 25 times instead of once. The average rate rose from 57 percent to 80 percent.

The attacker’s goals. Some goals are easier to hit, so an average depends on which goals are in the set. Across NIST’s five, some rates sat well above the 57 percent average and some markedly below.

The scorer. AgentDojo checks the state of its simulated environment; jailbreaking benchmarks such as HarmBench judge the model’s reply. Scorers do not always agree.

Why it matters

A low attack success rate does not mean a real attacker will usually fail. It reflects one set of attacks, one attempt budget, one set of tasks and one scorer.

In The Attacker Moves Second (2025), Milad Nasr, Nicholas Carlini and colleagues attacked 12 recent defences against jailbreaks and prompt injection. Most of those defences had reported near-zero rates. The authors bypassed all 12, most at above 90 percent.

A low rate still matters when the harm is large. NIST contrasts a harmless email with mass data theft, and warns that a low rate on its data exfiltration task “doesn’t mean this scenario should not be seriously considered and mitigated against.”

Where definitions disagree

“Success” depends on how it is measured. Earlier jailbreak papers often scored a reply by checking it for fixed strings of text. HarmBench found that, under those scorers, the length of reply the model was allowed to generate could change the rate by up to 30 percent. Because papers had not standardised that length, HarmBench called cross-paper comparisons “effectively meaningless”. The StrongREJECT authors found that existing scorers “significantly overstate jailbreak effectiveness” compared with human judgment.

Success can also mean different outcomes. AgentDojo’s targeted rate counts only cases where the attacker’s own goal was met. It also reports an untargeted rate: the share of cases where the agent did not finish the user’s task correctly or left side effects of the attack. That count includes tasks the agent might have failed with no attack at all.

“Attack” is not fixed either. AgentDojo sometimes scores a collection of attacks and counts a test case as broken if any one of them succeeds, which models an attacker who picks the best attack for each case. HarmBench’s rate is per attack method. Two papers can report an ASR for the same model and be measuring different things.

Questions and answers

What is a good attack success rate?

No attack success rate is good on its own, because the number describes the test as well as the system. NIST, working with the UK AI Security Institute, tested an agent built on Anthropic's upgraded Claude 3.5 Sonnet. On tasks held back from attack development, the strongest existing attack succeeded 11 percent of the time and the strongest new attack 81 percent. A rate is only meaningful next to the attacks, attempt budget, tasks and scorer that produced it.

Does a low attack success rate mean an AI system is secure?

No. A low attack success rate says the attacks in that test mostly failed. Researchers who built attacks tailored to each of 12 recently proposed defences for language models bypassed all 12, most at above 90 percent, though most of those defences had originally reported near-zero rates.

Sources

  1. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsarXiv, 19 Jun 2024
  2. Technical Blog: Strengthening AI Agent Hijacking EvaluationsNIST Center for AI Standards and Innovation, 17 Jan 2025
  3. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalarXiv, 6 Feb 2024
  4. A StrongREJECT for Empty JailbreaksarXiv, 15 Feb 2024
  5. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsarXiv, 10 Oct 2025

In the news