What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

STAC (sequential tool attack chaining)

STAC, short for Sequential Tool Attack Chaining, is a multi-turn attack on tool-using AI agents in which each tool call looks harmless alone, but the sequence together achieves a harmful goal only visible at the final step. It exploits agents that judge each call on its own, naming a method that generated and executed 483 such chains.

Last reviewed

Key points

  • STAC chains tool calls that each pass a safety check alone; only the final call reveals the harmful goal, after the earlier steps already set it up.
  • The coining paper generated and attacked with 483 STAC cases, reporting 91.2% average attack success across eight LLM agents — over 90% for seven of eight.
  • Per-call defenses barely notice it: prompt harmfulness stayed under 2% and refusal rate under 4%, because each request genuinely looks benign alone.
  • It beat a single-turn version of the same attack (72.8%) and two other multi-turn jailbreaks, MT-AgentRisk (53.2%) and X-Teaming (61.5%), in fewer turns.
  • The authors' reasoning defense cut first-turn success by 28.8 points, but an adaptive attacker eroded it within two turns; no tested defense held success below 74%.

STAC (sequential tool attack chaining) takes an AI agent through a sequence of tool calls that each pass a safety check on their own. Compress a file, delete the original, clean up leftover archives — three ordinary requests. Only the third one, run after the first two, destroys the only copy of a document that mattered. No single step in that sequence would have been flagged.

How it works

An automated pipeline builds each attack in four stages. A Generator plans a chain of two to six target tool calls ending in a malicious goal — say, deleting a file with no other copy — without touching a real system. A Verifier executes each planned call against the real environment and revises anything that fails, so the chain genuinely works. A Prompt Writer reverse-engineers user requests that would plausibly lead to each verified call, building a synthetic conversation up to the second-to-last step. A Planner then steers a live target agent through the final turns needed to trigger the harmful call.

The constraint enforced throughout is that every step but the last must be individually benign — so there is no malicious request for a per-request check to catch until the sequence is nearly done.

Why it matters

Most safety mechanisms in tool-using agents check the current tool call: does this specific request look dangerous? STAC is built to defeat exactly that check. Prompt harmfulness stayed under 2% and refusal rate stayed under 4% across every one of the eight agents the paper tested, because the individual requests genuinely do look ordinary. A per-call filter has nothing to catch.

That matters more for agents than chatbots. The paper distinguishes tool-mediated environmental change from harmful text generation: a jailbroken chatbot produces bad output a reader can ignore, but an agent that deletes a file or moves money has changed the world, sometimes irreversibly, before anyone reads the transcript.

In practice

The chain outperforms the single request. The paper’s clearest evidence that sequencing itself does the work is its own ablation: removing the synthetic multi-turn setup and asking for the same final malicious action in one prompt still worked 72.8% of the time, but the full multi-turn chain worked 95.1% of the time, in fewer turns. The gap is the contribution of spreading intent across verified, benign-looking steps rather than stating it once.

Defenses buy a little time, not safety. A reasoning-based defense that has the agent weigh harms and benefits before acting cut first-turn success sharply, but an attacker that kept adapting across two more turns eroded almost all of that gain within the same conversation. An experience-based defense that learned from prior attack trajectories held up better across turns but still left final attack success above three in four cases. Across every defense and every model family the authors tried, no combination brought success below 74%.

It is not limited to a lab simulation. Applying the same pipeline to live Model Context Protocol tool backends — a filesystem server and a browser-automation server — the attack still succeeded more than a third of the time in a fully black-box setting, where the attacker could see only tool schemas and had to discover the environment through its own tool calls, rather than the full environment visibility it had in simulation.

Trade-offs

The paper is explicit about the limits of its own evidence. Attack success is scored by a GPT-4.1 Judge that shares a model family with the Generator and Planner used to build the attacks, which the authors flag as a single-model-family dependency that could shape the reported numbers; they list cross-family judging and human validation as important follow-up work, not something already done. The headline results are also single-run point estimates without confidence intervals, and the released benchmark keeps only chains that passed in-environment verification, which the paper notes may bias the set toward more readily executable attacks. The real-tool MCP evaluation, while it confirms the attack transfers out of simulation, covers only two tool backends and a single agent, GPT-4.1.

Questions and answers

What does STAC stand for?

STAC stands for Sequential Tool Attack Chaining. It names a multi-turn attack in which an AI agent is led through a sequence of tool calls that each look harmless on their own, but that together accomplish a harmful goal that only becomes visible at the final step.

Is STAC a prompt injection attack?

No. The paper's threat model has the attacker issuing prompts directly through the normal user interface, not injecting content through a compromised document or tool output, and explicitly leaves indirect prompt injection to future work. STAC's stealth comes from distributing intent across multiple turns, not from hiding the attack inside data the agent reads.

How effective was STAC against real AI agents?

Across 483 generated attack cases against eight LLM agents including GPT-4.1, the paper reports a mean final attack success rate of 91.2%, exceeding 90% for seven of the eight agents tested. On a shared 390-case comparison, full multi-turn STAC reached 95.1%, against 72.8% for a single-turn version of the same attack and 53.2-61.5% for two other multi-turn jailbreak methods.

Does any defense stop STAC?

Not reliably. The paper's own harm-benefit reasoning prompt cut first-turn attack success by 28.8 percentage points on GPT-4.1, but an attacker that kept adapting over two more turns pushed success back up to 86.7%. An experience-based defense, ToolShield, held up best at 78.1% final attack success, still far from safe. Across every defense and every model family the authors tested, final attack success never fell below 74%.

Sources

  1. STAC: When Innocent Tools Form Dangerous Chains for LLM AgentsarXiv, 30 Sep 2025
  2. feat(attacks): STAC — sequential tool attack chaining (multi-turn), hackagent issue #600AISecurityLab (hackagent)

Guides that use this term