What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Adversarial suffix

An adversarial suffix is a string of tokens appended to a harmful request so that an aligned language model answers instead of refusing. The string is found by automated search, not written by hand. The best-known method, Greedy Coordinate Gradient (GCG), was published in 2023; its suffixes, built on open models, also worked on some closed ones.

Last reviewed

Key points

  • An adversarial suffix leaves the harmful request unchanged and adds a string after it, chosen by search to make the model start its answer with "Sure, here is". Nothing in the search rewards readable text, so the string often reads as gibberish.
  • The search needs a model's gradients, so it runs on open models. GCG tries batches of single-token swaps and keeps the best of each batch, step after step.
  • One suffix can be universal and transferable. A single suffix optimised on open Vicuna and Guanaco models raised GPT-3.5's rate of harmful answers from 1.8% to 47.4% across 388 requests; Claude 2 went from 0% to 1.8%.
  • Gibberish suffixes are easy to detect. In a 2023 test, perplexity filters caught nearly every GCG suffix but also flagged about one in ten ordinary prompts. A later method, AutoDAN, produces readable suffixes that pass.

How it works

The request stays as written and a string of tokens goes after it: 20 in Zou and colleagues’ 2023 experiments. For a harmful request, the search aims at the answer’s opening: “Sure, here is how to build a bomb”. Once a model starts that way, the authors found, it tends to carry on.

Greedy Coordinate Gradient (GCG) does the searching. Tokens cannot be nudged smoothly like the pixels of an adversarial example image, so GCG uses the model’s gradients to shortlist replacement tokens at every position, tries a batch of single swaps, and keeps whichever makes the affirmative opening most likely, over hundreds of steps.

Two choices make the result reusable. The suffix is optimised against many harmful requests, so it works on unseen ones, and against several open models, which helps it transfer to closed models whose gradients the attacker cannot see.

Why it matters

GCG beat earlier automated searches where they were weak. On Llama-2-7B-Chat, a suffix optimised on 25 harmful requests worked on 84% of 100 unseen ones, against 35% for AutoPrompt, the closest earlier method. On Vicuna-7B both reached 98%.

Unlike hand-written jailbreaking, GCG needs compute and an open model rather than a person’s ingenuity. Most model alignment training, the authors note, targets hand-written attacks, and they suspect automated search may make much of it insufficient.

Transfer to closed models was uneven. Across 388 harmful requests, one suffix raised the harmful-answer rate on GPT-3.5 from 1.8% to 47.4%. Claude 2 moved from 0% to 1.8%. Vicuna was trained on ChatGPT output, which the authors say may explain part of the GPT-3.5 result.

The authors warned OpenAI, Google, Meta and Anthropic first and expected their exact suffixes to stop working. Whether the underlying problem can be fixed at all, they wrote, was unclear.

Trade-offs

The obvious defence is to flag unnatural text. GCG suffixes are usually gibberish, so a perplexity filter, which scores how surprising a prompt is to a language model, can catch them. Jain and colleagues set filter thresholds so no plain harmful request was flagged. In their 2023 tests on five 7B-class models, the filters then caught all or nearly all GCG suffixes. They also flagged about one in ten ordinary prompts, which the authors called untenable on its own. Adding a low-perplexity goal to GCG’s search dragged its success down toward that of no attack at all.

That result held for GCG, not for every method. Later in 2023, AutoDAN optimised for readability as well as success and produced readable suffixes that passed perplexity filters.

The attack is also expensive. Jain and colleagues estimate GCG at 5 to 6 orders of magnitude more costly than attacks on image models, which makes the attacker’s compute budget a real limit.

Questions and answers

What does GCG stand for in AI security?

GCG stands for Greedy Coordinate Gradient, the search method Zou and colleagues published in 2023 for finding adversarial suffixes. At each step it uses the model's gradients to shortlist promising token swaps at every position in the suffix, tests a batch of them, and keeps the one that most raises the chance the model starts its answer with "Sure, here is".

Does an adversarial suffix still work on ChatGPT or Claude?

The 2023 paper cannot say, and this page does not track today's models. The authors shared their results with OpenAI, Google, Meta and Anthropic before publishing and expected the paper's examples to stop working. The attack code is public, and the authors said it was unclear whether the underlying problem could be fixed at all.

How is an adversarial suffix different from an ordinary jailbreak?

An ordinary jailbreak is typically crafted, in Zou and colleagues' words, through "human ingenuity". An adversarial suffix is found by an optimiser using a model's gradients. It takes compute rather than ingenuity, often reads as gibberish, and so is easier to spot with a filter that flags unnatural text.

Sources

  1. Universal and Transferable Adversarial Attacks on Aligned Language ModelsAndy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson (arXiv), 27 Jul 2023
  2. Baseline Defenses for Adversarial Attacks Against Aligned Language ModelsNeel Jain et al., University of Maryland (arXiv), 1 Sep 2023
  3. AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language ModelsSicheng Zhu et al., University of Maryland and Adobe Research (arXiv), 23 Oct 2023

Guides that use this term