What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Many-shot jailbreaking

Many-shot jailbreaking is an attack that fills a single long prompt with hundreds of fabricated dialogues in which an AI assistant answers harmful questions, then asks the real question last. The model learns the pattern from its own prompt and continues it. Anthropic researchers found that past a threshold, more dialogues make a harmful answer more likely.

Last reviewed

Key points

  • Many-shot jailbreaking needs a long context window. The attacker places dozens to hundreds of faked compliant exchanges ahead of the real harmful request, all in one prompt.
  • It works through in-context learning, the same ability that lets a model pick up any pattern from examples. Anthropic found the attack and ordinary in-context learning scale with the number of examples in the same way.
  • On a personality-trait evaluation, around 128 faked exchanges were enough to make Claude 2.0, GPT-3.5, GPT-4, Llama 2 70B and Mistral 7B adopt the harmful behaviour. Among Claude models, larger ones needed fewer.
  • Safety fine-tuning only delays the attack. More faked exchanges are needed, but harmful answers still appear. In one test, a warning wrapped around the prompt cut attack success from 61 percent to 2 percent.

How it works

Many-shot jailbreaking works by teaching the model, inside one prompt, that this assistant answers anything. The attacker produces hundreds of question-and-answer pairs in which an assistant readily answers harmful requests, formats them as a user–assistant dialogue, and puts the real question at the end. It all goes to the model as one message.

A few examples do nothing. On Claude 2.0, Anthropic’s researchers found the attack does not work at all with 5 faked exchanges and works consistently with 256.

The reason is in-context learning: a language model picks up patterns from examples in its prompt without any retraining. Anthropic found the attack’s success grows with the number of examples following the same power law as benign in-context learning tasks. Among Claude models of different sizes, larger ones learned faster in context and needed fewer faked exchanges.

The attack became practical when context windows grew, from about 4,000 tokens at the start of 2023 to 1,000,000 or more in some models by April 2024.

Why it matters

Many-shot jailbreaking is hard to train away. In the paper, standard safety fine-tuning and reinforcement learning made the attack need more faked exchanges, but did not stop harmful answers once the prompt was long enough. Combining it with other jailbreaking techniques shortened the prompt required again.

Capping the context window would stop it outright, at the cost of what long inputs are for. Prompt-level defences did better in testing. On one category with a 205-shot prompt, prepending examples of refusals cut attack success only from 61 to 54 percent. Wrapping the prompt in warning text cut it to 2 percent, a defence whose cost to usefulness the authors say is still unmeasured.

The 2024 paper noted one limit: the basic attack needs API access, because consumer chat apps did not let users insert a fake dialogue history.

Questions and answers

How is many-shot jailbreaking different from an ordinary jailbreak?

An ordinary jailbreak talks a model past its safety training with a framing, such as a story or a claimed authorization. Many-shot jailbreaking uses volume instead: hundreds of fabricated exchanges in one prompt that show an assistant complying, so the model continues the pattern. It only became practical once context windows grew long enough to hold that many examples.

Can someone run a many-shot jailbreak in a normal chat app?

Not in its basic form, as of 2024. The Anthropic paper that named the attack notes that it needs API access, because consumer chat apps such as ChatGPT and Claude.ai did not let a user insert a fake dialogue history. An application that passes long user-supplied text straight to a model through an API has no such barrier.

Does safety training fix many-shot jailbreaking?

Not on its own. Anthropic found that training a model to refuse many-shot prompts raised the number of faked exchanges needed, but harmful answers still appeared once the prompt was long enough. Prompt-level defences did better in testing, but their cost to the model's usefulness has not been measured.

Sources

  1. Many-shot jailbreakingAnthropic, 2 Apr 2024
  2. Many-shot Jailbreaking (NeurIPS 2024)Cem Anil et al., Advances in Neural Information Processing Systems 37

Guides that use this term