Guide · AI security
Guardrail bypass techniques
A language model can be protected by several controls: its safety training, guardrails that check what goes in and comes out, and a trained ranking of whose instructions win. Bypass techniques are grouped here by the control each one mainly gets past, followed by the reconnaissance that shows an attacker which rules are in place.
Last reviewed
The model’s own safety training
The request reaches the model, and the model is persuaded to answer what it was trained to refuse. See model alignment for what these undo.
- Jailbreaking: reframing a request as fiction, a hypothetical or an authorised test so the model sets its safety training aside.
- Many-shot jailbreaking: hundreds of faked compliant exchanges in one long prompt, so the model continues the pattern on the real question.
- Adversarial suffix: a string of tokens found by automated search and appended to a harmful request, so the model begins its answer with compliance instead of a refusal. Suffixes built on open models also worked on some closed ones.
- Salami Attack: a chain of prompts that each stay below the refusal threshold, adding up to a request the model would have refused outright. Each prompt also looks harmless to a filter that judges prompts one at a time.
- Crescendo: a conversation that opens with a harmless question and escalates step by step, each prompt building on the model’s own last answer instead of stating the goal.
The filter in front of the model
The text is written so a checker reads it as harmless while the model behind it reads the real instruction. See AI guardrails for the checks these target.
- ASCII smuggling: instructions in invisible Unicode tag characters, which a human reviewer cannot see and filters that do not strip them miss, while a model that decodes them reads the instruction.
- PuzzleMask: a forbidden instruction hidden in ordinary prose. In Check Point’s tests, four screening models passed every crafted prompt as safe, and a reasoning model with a code interpreter behind them decoded and acted on it.
The trust given to where text comes from
The instruction arrives where the model expects only content to process, or carrying a role it did not earn. See Instruction hierarchy for the ranking these get around.
- Indirect prompt injection: instructions hidden in a page, email or file the model later reads, which the user never sees.
- Structured template injection: forged chat-template turn markers in retrieved content, so attacker text reads as a user request.
- Role confusion: untrusted text written to sound like a trusted speaker, so the model may take it as that speaker even though it arrived labelled as something else.
- Prompt Forcing: a browser extension or website writes the agent’s prompts itself, and the agent treats them as coming from the user.
- Instruction privilege escalation: attacker text that entered as a tool result is passed on, for example as a delegated task, and the agent harness writes it into a user message.
Removing the safety training
When the attacker can change the model, no prompt is needed: the refusals are trained or edited out.
- Fine-tuning: further training that can undo safety training, through a fine-tuning API as well as on downloaded weights.
- Abliterated vs uncensored models: abliteration edits out the direction in a model’s activations that carries refusal, with no retraining.
Learning the rules first
Reconnaissance that gets past no control by itself, but shows the attacker what to aim at.
- System prompt extraction: recovering a model’s hidden instructions, which shows an attacker which rules to route around.