What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Salami Attack

Salami Attack is an automated multi-turn jailbreak framework that chains a harmless opening prompt with several harmless-seeming follow-up prompts, each individually below a target model's refusal threshold, so their combined effect accumulates into a harmful output the model would have refused if asked directly. It exploits what its authors call "Salami Slicing Risk."

Last reviewed

Key points

  • Salami Attack exploits "Salami Slicing Risk," chaining low-risk inputs that each individually evade a model's alignment threshold but cumulatively accumulate harmful intent, without needing a pre-scripted scenario or a tool call.
  • An LLM automatically drafts one harmless opening prompt plus several harmless-seeming perturbations, repeats them to amplify their effect, then feeds the whole sequence to the target model.
  • The paper reports over 90% attack success rate on GPT-4o and Gemini, and calls the framework "universally applicable to multiple model types and modalities", extending it to vision-language and diffusion models too.
  • The authors' own defense, Cumulative Query Auditing, re-scores a model's entire conversation history after every new prompt rather than just the newest one, but still only cuts the attack's success rate by 44.8% on GPT-4o.

How it works

Most models refuse a request that is obviously harmful, and multi-turn jailbreaking has long tried to get around that by splitting the request into several “seemingly safe steps”. The paper behind Salami Attack argues that approach still fails often, because an explicit harmful trigger “is increasingly likely to be flagged and blocked” and because making the final step actually land usually takes a hand-tuned scenario built for that one model.

Salami Attack automates a different split. An LLM first drafts one opening prompt that matches the topic and tone of the eventual harmful request but stays harmless on its own. It then drafts several more prompts, each nudging the conversation slightly further, with none crossing the line a safety filter would catch alone. The full sequence, each prompt repeated to strengthen its effect, is fed to the target model in order — by the last prompt, the model’s own accumulated context does the work a blunt request never could.

Why it matters

The paper reports over 90% attack success rate on GPT-4o and Gemini, with per-model AdvBench results of 90.4% (GPT-4o), 91.5% (Gemini 2.5 Pro), 96.7% (DeepSeek V3), 97.7% (Qwen 3) and 93.3% (GLM-4.5) — the same automated pipeline generating the chain each time, which is what the authors mean by calling it “universally applicable to multiple model types and modalities.”

That reach isn’t limited to chat. The authors extend the same principle to vision-language models, where interleaved images and text give a filter more surface to miss cumulative intent across, and to diffusion models, where a chain of small, innocuous prompt edits incrementally steers image generation toward output a single explicit prompt would have been refused for.

In practice

A safety system built to judge each incoming prompt on its own has nothing to catch here — every individual prompt in the chain is, by design, below that per-message bar. Catching Salami Attack means judging the conversation, not the message.

That’s the shape of the authors’ own proposed defense, Cumulative Query Auditing (CQA): after every new prompt, it reassembles the entire conversation so far into one sequence and scores that cumulative sequence for harm, refusing once the running total crosses a threshold, rather than scoring only the newest message. On GPT-4o, adding CQA cut the attack’s success rate from 87.0% to 48.0%, a 44.8% reduction the authors call their headline result, even though the paper also reports CQA cutting a comparison attack from 54.0% to 19.0%, a 64.8% reduction — doing less well against the attack it was built for than against another one.

Trade-offs

Cutting a 5-shot Salami Attack’s success rate to 48.0% still means it succeeds close to half the time against the defense purpose-built to stop it. The paper does not report a case where any tested defense drove the attack’s success meaningfully below that. For a reader deciding how much weight to put on a defense here, the honest number is the 44.8% reduction, not the framing of it as a solved problem.

Salami Attack is often the nearer read to STAC, another chained, per-step-innocuous attack, but the two target different surfaces: STAC chains tool calls an agent makes against a real environment, while Salami Attack chains ordinary prompts a user sends in chat, with no agent or tool access required.

Questions and answers

What is a Salami Attack?

Salami Attack is an automated multi-turn jailbreak framework that chains a harmless-looking opening prompt with several more harmless-seeming follow-up prompts. No single prompt in the chain is dangerous enough to trigger a model's refusal, but their combined effect accumulates into a harmful output the model would have refused if it had been asked directly.

What is Salami Slicing Risk?

Salami Slicing Risk is the underlying vulnerability the paper that coined the term names — chaining numerous individually low-risk inputs that each evade a model's alignment threshold on their own, but cumulatively build up harmful intent, without relying on a pre-scripted scenario. Salami Attack is the automatic framework built to exploit it.

How effective is Salami Attack?

The paper reports over 90% attack success rate on GPT-4o and Gemini, with per-model results on AdvBench of 90.4% (GPT-4o), 91.5% (Gemini 2.5 Pro), 96.7% (DeepSeek V3), 97.7% (Qwen 3) and 93.3% (GLM-4.5) using a 10-shot chain of prompts.

Does Salami Attack only work on chatbots?

No. Its authors report it applies across "multiple model types and modalities", extending the same low-risk-chaining principle to vision-language models, which process interleaved images and text, and to diffusion models, where each small, incremental edit to a generated image stays below the safety threshold while the cumulative sequence steers the output toward disallowed content.

What stops Salami Attack?

The paper's own defense, Cumulative Query Auditing (CQA), re-audits the full conversation history after every new prompt instead of judging each prompt alone, and reduced the attack's success rate on GPT-4o from 87.0% to 48.0%, a 44.8% relative cut. That leaves it succeeding roughly half the time even against the defense built specifically to counter it.

How is Salami Attack different from STAC?

STAC chains agentic tool calls against a real environment, so it requires an agent that can act. Salami Attack chains ordinary conversational or multimodal prompts with no tool-use requirement, and its authors report it generalizing across model families and across text, vision-language, and diffusion modalities.

Sources

  1. The Salami Slicing Threat: Exploiting Cumulative Risks in LLM SystemsarXiv, 13 Apr 2026
  2. The Salami Slicing Threat: Exploiting Cumulative Risks in LLM SystemsarXiv, 13 Apr 2026

Guides that use this term