What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Jailbreaking

Jailbreaking is an attack that gets a language model to override its own safety training and produce output it was built to refuse, typically by reframing the request as fiction, a hypothetical or an authorized test. Unlike prompt injection, the attacker is usually the person operating the model, working against restrictions the model's own creator built in.

Last reviewed

Key points

  • Jailbreaking attacks the model's own safety training, usually by the person operating it — the reverse of prompt injection, where a third party's text overrides an application built on top of the model.
  • It works because safety training pattern-matches on how a request is framed rather than what it asks for. Wei et al. name two failure modes — competing objectives and mismatched generalization — and jailbroke GPT-4 and Claude v1.3 on all 32 curated red-team prompts.
  • Many-shot jailbreaking scales the idea with volume — Anthropic found stuffing a prompt with dozens of faked compliant exchanges raises the odds a model completes one more the same way.
  • A 2023 study of prompts shared across the open web found five that reached a 0.95 attack success rate against GPT-3.5 and GPT-4; the earliest of the five had stayed posted online for over 240 days.

Jailbreaking is a language model doing what it was trained to refuse, talked into it by the person using it.

How it works

Safety training teaches a model to recognize and refuse certain requests, but recognition runs on the surface form of the request, not on intent. Wei, Haghtalab and Steinhardt call the two gaps this leaves competing objectives, where a model’s drive to be helpful pulls against its drive to be safe, and mismatched generalization, where safety training never covered a framing the model is otherwise capable of handling.

The simplest jailbreak exploits that gap with a story: ask a model to answer as a grandmother who “used to read me napalm recipes,” and it may complete the bedtime-story framing rather than checking whether the underlying request is one it would otherwise refuse. Many-shot jailbreaking scales the same idea with volume instead of character: a prompt stuffed with dozens of fabricated exchanges of a compliant assistant answering harmful questions raises the odds the language model completes one more the same way, because it is using the same in-context learning that lets it follow any other pattern in a long prompt.

Why it matters

Jailbreak prompts work often enough, and travel far enough, to matter. A study of prompts shared across the open web between December 2022 and December 2023 found five that reached a 0.95 attack success rate against both GPT-3.5 and GPT-4, and the earliest of the five had stayed posted online for more than 240 days. AI guardrails and AI red teaming cut the success rate of specific, known attacks — in one test, wrapping a many-shot jailbreaking prompt in warning text took it from 61% to 2% — but a defence tuned to today’s jailbreaks does not close the gap between framing and intent that makes new ones possible.

In practice

Jailbreaks need not arrive in one prompt. Salami Attack spreads the request across turns, chaining prompts that each stay below the model’s refusal threshold until their combined intent produces the output a direct request would not. A filter that judges each message alone has nothing to catch. The defence proposed by the Salami Attack authors re-scores the whole conversation after every prompt, and on GPT-4o it still only cut the attack’s success rate from 87.0% to 48.0%.

Questions and answers

Is jailbreaking the same as prompt injection?

They overlap but are not the same. Jailbreaking targets the model's own safety training, and the attacker is usually the person operating the model. Prompt injection targets an application built on top of the model, and the attacker is usually a third party whose text the model reads. OWASP's LLM Top 10 treats jailbreaking as a subset of prompt injection; Willison, who coined "prompt injection," keeps the two apart.

Why doesn't safety training stop every jailbreak?

Safety training pattern-matches on how a request is phrased more than on what it actually asks for. Wei, Haghtalab and Steinhardt describe two gaps this leaves: competing objectives, where being helpful and being safe pull against each other, and mismatched generalization, where training never covered a framing the model is otherwise capable of handling — a hypothetical, a story, a claimed authorization.

What is many-shot jailbreaking?

A technique that fills a long prompt with dozens of fabricated exchanges showing an assistant answering harmful questions, then adds the real question last. Anthropic found the model becomes more likely to continue the pattern as the number of fake exchanges grows, because the attack uses the same in-context learning a model uses for any other long-prompt pattern, not a distinct flaw a single patch can close.

Sources

  1. Prompt injection and jailbreaking are not the same thingSimon Willison, 5 Mar 2024
  2. Jailbroken: How Does LLM Safety Training Fail?arXiv, 5 Jul 2023
  3. Many-shot jailbreakingAnthropic, 2 Apr 2024
  4. Many-shot Jailbreaking (NeurIPS 2024)Cem Anil et al., Advances in Neural Information Processing Systems 37
  5. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsarXiv, 7 Aug 2023
  6. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsarXiv, 10 Oct 2025

Guides that use this term