Definition · AI agents
LoopTrap
LoopTrap is a red-teaming framework that attacks termination poisoning — an adversary corrupting the progress signals an iterative LLM agent uses to judge whether a task is complete, so the agent believes it must keep going. It profiles a target agent's termination biases, then synthesizes and injects a tailored trap to keep it looping past actual completion.
Last reviewed
Key points
- LoopTrap targets iterative LLM agents, which decide for themselves when a task is finished; the paper names the attack surface it exploits termination poisoning.
- It profiles an agent's termination biases across four dimensions, then synthesizes an injection tailored to that profile, and banks successful attacks in a reusable skill library.
- Ten named attack strategies fall into four families — progress manipulation, cognitive bias exploitation, task structure manipulation, and reward shaping — such as telling the agent its 92%-complete estimate needs "just a bit more refinement."
- Across 8 mainstream agents and 60 GAIA-benchmark tasks, LoopTrap reached an average attack success rate of 86.3% and an average step-amplification factor of 3.57x (peak 25x), beating the strongest static baseline by 14.6 points of ASR.
- Agents differ in which strategy works — Kimi-K2-Thinking defers to a fabricated authority directive but resists recursive-decomposition traps, Claude Sonnet 4.5 shows the opposite pattern — and profiling alone lifts ASR from 81.7% to 86.3%.
How it works
Iterative LLM agents — the kind that plan, act, observe, and repeat until they judge a task finished — have no external referee for that last step. LoopTrap treats that self-judgment as the attack surface. It first sends an agent lightweight diagnostic probes to fingerprint its termination biases along four dimensions: how readily it defers to authority, how easily it’s strung along by phased sub-goals, how thoroughly it verifies before stopping, and how susceptible it is to recursive task decomposition.
Guided by that profile, adaptive trap synthesis picks the attack strategy most likely to work on this particular agent and task, generates candidate injections, and self-scores them to choose the one most likely to disrupt the agent’s progress assessment. A third stage abstracts successful attacks into a reusable skill library, so what works against one agent or task informs the next attempt.
Ten named strategies fall into four families: progress manipulation (a moving completion target, never-ending sub-goals, an asymptotic “you’re 92% done”), cognitive bias exploitation (fabricated authority directives, sunk-cost framing, social proof), task structure manipulation (recursive decomposition, dependency chains), and reward shaping (positive reinforcement, gamification).
Why it matters
The attack surface is the agent’s autonomy itself: nothing outside the agent stops it once it’s convinced there’s still work to do. Across 8 mainstream agents and 60 real-world GAIA-benchmark tasks, LoopTrap reached an average attack success rate of 86.3% and an average step-amplification factor of 3.57x, with a peak of 25x on individual tasks, while also increasing total token overhead by 3.93x — beating the strongest static baseline by 14.6 points of ASR.
In practice
Which strategy works depends on the agent, not just the attack. Kimi-K2-Thinking readily defers to a fabricated authority directive and abandons a near-correct reasoning path, but resists recursive-decomposition traps; Claude Sonnet 4.5 shows close to the opposite pattern, while Grok-4 falls into a prolonged loop under recursive decomposition. The profiling step itself matters: removing it drops ASR from 86.3% to 81.7% and the step-amplification factor from 3.57x to 2.95x.
The paper proposes two defensive directions without testing them to completion — an independent module that checks an agent’s self-reported progress against objective criteria, and context processing that weighs progress signals differently by whether they come from trusted instructions or untrusted retrieved content — but calls both preliminary.
Where definitions disagree
The term is new: LoopTrap and termination poisoning come from a single May 2026 arXiv submission by researchers at Zhejiang University and Southeast University, with no other source using either name yet. The paper is careful to place its target next to, not inside, prompt injection: ordinary prompt injection changes what a single generation says, while termination poisoning changes whether the agent’s control loop ever decides to stop. It’s also distinct from agent hijacking, which redirects an agent to attacker-chosen actions — termination poisoning leaves the agent doing on-task work, just for far longer than the task requires.
Questions and answers
What is LoopTrap?
LoopTrap is a red-teaming framework that attacks termination poisoning, the vulnerability in iterative LLM agents that must judge for themselves when a task is complete. It profiles a target agent's termination biases along four dimensions, then synthesizes and injects a trap tailored to that profile and task, keeping the agent looping past actual completion.
What is termination poisoning?
Termination poisoning is the paper's name for corrupting the progress signals an LLM agent uses to decide a task is done: "an adversary injects malicious content into the agent's operational context to corrupt the progress signals the agent uses to assess task completion, thereby preventing termination and inducing unbounded execution loops."
How effective is LoopTrap?
Across 8 mainstream agents and 60 real-world GAIA-benchmark tasks, LoopTrap reached an average attack success rate of 86.3% and an average step-amplification factor of 3.57x, with a peak of 25x on individual tasks — beating the strongest static baseline by 14.6 points of ASR.
How is termination poisoning different from prompt injection?
The paper draws the line at what's targeted. Conventional prompt injection changes the output of a single generation; termination poisoning targets the agent's control flow — the mechanism that decides whether to keep running — leaving individual outputs otherwise on-task.
Are there defenses against LoopTrap?
The paper proposes two untested directions — an independent module that verifies an agent's self-assessed progress against objective criteria, and provenance-aware context processing that weighs progress signals differently depending on whether they come from trusted instructions or untrusted retrieved content — but calls both "preliminary suggestions" with limited empirical validation.
Sources
- LoopTrap: Termination Poisoning Attacks on LLM AgentsarXiv, 7 May 2026
- LoopTrap: Termination Poisoning Attacks on LLM AgentsarXiv, 7 May 2026