Definition · AI security
LLM prompt self-replication
LLM prompt self-replication is an attack in which a crafted prompt injection makes a large language model reproduce the malicious prompt in its own output. The copy persists and spreads to other models that read that output. The prompt usually also carries other malicious instructions, such as a jailbreak or data leakage. MITRE ATLAS catalogues the technique as AML.T0061.
Last reviewed
Key points
- LLM prompt self-replication is a prompt injection that tells the model to copy the injected prompt into its output. Wherever that output is stored or sent, the prompt goes with it. MITRE ATLAS catalogues it as AML.T0061, a persistence technique.
- The copied prompt usually carries other malicious instructions, such as a jailbreak that makes the model obey or an order to leak data. Each system that copies the prompt can act on them too.
- ATLAS's one recorded case is Morris II, a 2024 lab worm that spread between RAG-based email assistants and exposed private data in the emails they drafted.
- ATLAS maps three defences that can help — guardrails, guidelines and model alignment. The Morris II authors built a guardrail that flags output copying its input; set to catch every worm, it raised far more false alarms in most tests on worms and emails unlike its training data, and it was not tested against an attacker adapting to it.
LLM prompt self-replication turns a single prompt injection into something that spreads. MITRE ATLAS files it as AML.T0061: a prompt injection “designed to cause the LLM to replicate the prompt as part of its output”, which “allows the prompt to propagate to other LLMs and persist on the system”.
How it works
The prompt has three parts, in the Morris II researchers’ description:
- A jailbreak, which they call “privilege escalation”: it gets the model to follow the rest.
- A copy instruction, telling the model, in the paper’s words, to “replicate the input into the output”.
- Malicious instructions, such as collecting private data from the context.
The model’s output then contains the whole prompt again. If that output is stored where another assistant will retrieve it, such as the inbox a retrieval augmented generation email assistant draws on, that assistant can be infected in turn. In Morris II the attacker only sent the first message; the researchers write that the attacker’s role “ends once the worm is launched”.
Why it matters
An ordinary prompt injection acts where the attacker places it. A self-replicating one writes itself into new places, so it can keep spreading without the attacker. ATLAS lists it under the Persistence tactic, and notes it is “typically paired with other malicious instructions”, such as a jailbreak or data leakage.
ATLAS maps three defences, each worded as help rather than a fix. AI guardrails “can help prevent replication attacks in model inputs and outputs”. Guidelines, such as a system prompt, can instruct the model not to produce self-replicating output. Model alignment “can increase the security of models” against these attacks.
In practice
ATLAS’s one recorded case is Morris II, a lab worm published in 2024 by researchers at Technion and Intuit. They emailed a self-replicating prompt to a RAG-based email assistant. The assistant stored it and, when a later request retrieved it, copied it into the emails it drafted, along with private data taken from the user’s other emails.
The January 2025 revision tested the worm against Gemini 1.5 Flash, with 20 emails placed in the model’s context and a worm that opened with text about the company whose mail it targeted. When replying to or expanding an email, the prompt was retrieved, copied and carried its payload about 20% of the time. The authors read that as the worm reaching a new user roughly once every five emails.
The authors also built a guardrail, the Virtual Donkey, which flags output that closely copies its input. Tuned to catch every worm in their main test, it raised a false alarm on about one clean email in 66. On worms and emails unlike its training data, catching every worm cost far more false alarms in most tests. The authors still call it robust out of distribution, on a different measure. They did not test it against an attacker who adapts to it.
Where definitions disagree
ATLAS treats self-replication as its own technique, under Persistence. OWASP’s 2026 LLM01 does not: it lists “self-replicating across agents” as one propagation behaviour of prompt injection, next to single-shot and cross-session. The Morris II paper names the result a worm, after the 1988 Morris worm. All three describe the same mechanism and file it differently.
Questions and answers
What is a self-replicating prompt?
A self-replicating prompt is a prompt injection that instructs a language model to reproduce the prompt in its own output. When that output is stored or sent to another AI system, the prompt travels with it and can act again. MITRE ATLAS catalogues the technique as AML.T0061, LLM Prompt Self-Replication.
What was the Morris II worm?
Morris II was a research worm built by Stav Cohen, Ron Bitton and Ben Nassi and published in 2024. An email carrying a self-replicating prompt was stored by a RAG-based email assistant. When the assistant later retrieved it, it copied the prompt, and private data from the user's other emails, into the emails it drafted, exposing that data to other users. It was run only in a lab.
How do you defend against self-replicating prompts?
MITRE ATLAS maps three mitigations that can help — guardrails on model inputs and outputs, guidelines such as system prompts, and model alignment. The Morris II researchers built a guardrail that flags output closely copying its input. Set to catch every worm, it raised far more false alarms in most tests on worms and emails unlike its training data, and it was not tested against attackers adapting to it.
Sources
- MITRE ATLAS, AML.T0061 LLM Prompt Self-Replication (collection 2026.09)MITRE
- Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications (arXiv:2403.02817v2)Stav Cohen, Ron Bitton and Ben Nassi (Technion, Intuit), 30 Jan 2025
- LLM01:2026 Prompt Injection, OWASP Top 10 for LLM ApplicationsOWASP