Definition · AI agents
Agent memory poisoning
Agent memory poisoning is an attack that writes attacker-controlled content into an AI agent's persistent memory, so the agent retrieves it as its own trusted context in later sessions. One successful write outlives the session that made it, and the agent acts on the planted content long after the attacker has gone.
Last reviewed
Key points
- Agent memory poisoning needs one successful write. Prompt injection needs the hostile text in the context window right now; a poisoned memory is read back as trusted context in sessions the attacker never touches.
- MITRE ATLAS files it as persistence, not execution. AI Agent Context Poisoning (AML.T0080) sits under the Persistence tactic, while LLM Prompt Injection (AML.T0051) sits under Execution.
- The write channel is the ordinary one. ATLAS notes that memory is controlled through normal conversation, so an adversary injects memories by direct or indirect prompt injection.
- Relevance decides whether a planted memory fires. In one 2026 study, agents acted on a retrieved poisoned memory in 60 to 89 percent of related agentic evaluations.
- Prompt injection defences do not cover it. They look for text that reads like an instruction; a poisoned memory need not, and write-time filters still pass records that only turn harmful once combined or triggered.
Memory poisoning is what makes a prompt injection outlast the conversation it arrived in. MITRE ATLAS files AI Agent Context Poisoning as AML.T0080 and puts it under Persistence, the tactic for keeping a foothold “across restarts, changed credentials, and other interruptions that could cut off their access”. Prompt injection itself sits under Execution.
How it works
Three steps, and the attacker is present for only the first.
The write. Attacker-controlled text reaches the agent through an ordinary channel — a shared document, a web page, a message — and the agent saves it. Nothing is exploited. ATLAS notes that memory is controlled through normal conversations with the user, so an adversary can inject memories by direct or indirect prompt injection.
The gap. Nothing happens, for as long as the attacker likes. No further access is needed.
The read. A later session retrieves the entry as the agent’s own context. The ATLAS sub-technique for memory, AML.T0080.000, describes manipulating an LLM’s memory “in order to persist changes to the LLM across future chat sessions”.
Why it matters
A poisoned memory is trusted in a way an incoming prompt is not. Microsoft’s attack catalogue says so directly: memory poisoning works because it exploits “trusted internal state”, not external prompts, and many systems “treat memory as authoritative”, ranking it above live input. It also notes that users and admins rarely see raw memory entries, so the corruption is subtle.
The defences already in place do not cover it, and not because they are looking in the wrong place. A 2026 systematic study concludes that “existing prompt injection defenses fail to cover memory poisoning attacks” — both arrive through inputs the agent processes in context, but those defences hunt explicit malicious instructions, and fail where “the payload carries no detectable pattern in the context”.
In practice
It has been demonstrated against a shipping product. ATLAS records case study AML.CS0040 against ChatGPT, typed as an Exercise rather than a real-world incident. A researcher hid an injection in a shared Google Doc; when a user referenced the document, ChatGPT wrote false facts to memory, and the chat window reported the memory set “despite the lack of human verification or intervention”. Every later session used the poisoned store.
A planted memory fires hardest when it is relevant. The sleeper memory poisoning study found that both retrieval and downstream use are strongly dependent on how close the poisoned memory is to what the user later asks. On its agent-action tests, once a goal-adjacent poisoned memory was retrieved, models acted on it in 60 percent of evaluations for Claude-Sonnet-4.6 and 89 percent for Gemini-3.1. That figure is conditional on retrieval happening at all. Composed end to end, with injection and retrieval folded back in, the same agent-action tests range from 3.0 to 66.0 percent across models — the paper’s 41.0 to 73.9 percent is its other subset, plain assistant behaviour rather than agent actions. A useful poisoned memory has to look like ordinary task context, which is also what makes it hard to spot.
Filtering each write on its own has a structural blind spot. MemPoison, a benchmark of 1,227 hand-validated cases, reports that write-time consistency checks substantially suppress direct single-record corruption but do not reliably suppress corruption spread across several records, or records that stay dormant until a particular context triggers them. Each record passes inspection; the harm appears when they are retrieved together.
More eager memory means more exposure. The same systematic study reports that agents designed to write and retrieve memory more aggressively are more exploitable — a design dial, not a bug to patch.
The only mitigation ATLAS scopes specifically to this technique is AML.M0031, Memory Hardening, which it distinguishes from guardrails on the grounds that guardrails evaluate content during an interaction while memory hardening controls how durable agent state is created, modified, isolated, audited and recovered. ATLAS names that guardrail control, AML.M0020, Generative AI Guardrails, and says content filtering and protections against malicious prompts “does not replace memory access control, isolation, provenance, integrity, or recovery”; in the pinned collection M0020 mitigates 23 techniques, including AML.T0051, LLM Prompt Injection, and none of the AML.T0080 family. Memory Hardening’s four named controls are memory access controls; memory security policies covering size, update frequency, integrity and retention; memory provenance — recording each update’s source, preserving known good versions, and quarantining or rolling back suspicious records; and auditing and monitoring of memory operations — auditing security-relevant reads and writes, and monitoring for unusual update frequency and repeated self-authored updates. It is not the only mitigation ATLAS maps here: AML.M0035, AI Red Team, covers the technique too, but as one of 33 targets spread across the matrix rather than a memory-specific control — and it is the only one of the two that reaches AML.T0080.001, Thread, which Memory Hardening does not mitigate.
Where definitions disagree
How much counts as memory. ATLAS splits the store from the conversation. Memory, AML.T0080.000, is a user-specific database, and poisoning it persists changes across future chat sessions; Thread, AML.T0080.001, contaminates the thread itself, for the remainder of that thread — which is not necessarily shorter, because ATLAS adds that a chat thread “may continue for an extended period over multiple sessions”. Both sit under Persistence. OWASP merges them: its ASI06, Memory & Context Poisoning, covers “any information an agent retains, retrieves, or reuses, such as summaries, embeddings, and RAG stores”, and draws its line at the prompt instead, excluding one-time input prompts as prompt injection. Under OWASP a poisoned retrieval index is the same risk; under ATLAS it is a different technique.
Whether the attack has a settled name at all. NIST’s adversarial machine learning taxonomy, AI 100-2e2025, published March 2025, does not use the term. Its section 3.5, Security of Agents, names memory only as a feature an agent may have, alongside planning, and says that “Security research focused specifically on agents is still in its early stages”, while noting that researchers have begun to evaluate agents against particular attacks. The term is carried by ATLAS, OWASP, vendor catalogues and a 2026 wave of preprints, not by NIST’s taxonomy, so a reader who goes looking for it there will not find it.
Questions and answers
Is memory poisoning just prompt injection?
No, though prompt injection is usually how the content gets in. The difference is where the attacker's text has to be at the moment it does damage. A prompt injection has to be in the context window the model is reading right now, and MITRE ATLAS notes that its effects can persist throughout an interactive session. A poisoned memory is read back out of the agent's own store in a later session, which is why ATLAS classifies it under Persistence rather than Execution.
How long does a poisoned memory last?
Until someone finds and removes it. MITRE ATLAS describes two forms. Thread poisoning, AML.T0080.001, causes behaviour changes which persist for the remainder of the thread, and ATLAS notes a thread may itself run over multiple sessions. Memory poisoning, AML.T0080.000, persists changes across future chat sessions, in a user-specific store. In the ATLAS case study of a demonstration against ChatGPT's memory feature, the procedure records that all future chat sessions will use the poisoned memory store.
Can poisoned memories be filtered out when they are written?
Partly, and the gap is structural. The MemPoison benchmark reports that baseline write-time defences such as consistency checks substantially suppress direct single-record corruption, but fail to reliably suppress corruption spread across several records or triggered later by a particular context. Those records look benign one at a time and only become harmful when retrieved together or activated by a trigger, so inspecting each write on its own cannot catch them.
Does memory poisoning need the attacker to keep access?
No, and that is the point of the attack. The write happens once, through an ordinary channel such as a shared document or a web page the agent reads during a normal task. After that the attacker needs no further access, because the agent itself retrieves the planted content and acts on it. This is why MITRE ATLAS places it under Persistence, the tactic for keeping a foothold across interruptions that would otherwise cut an adversary off.
Sources
- MITRE ATLAS, technique AML.T0080 AI Agent Context Poisoning (collection 2026.08)MITRE
- OWASP Top 10 for Agentic Applications 2026, ASI06: Memory & Context PoisoningOWASP Gen AI Security Project, 9 Dec 2025
- AI memory/context poisoningMicrosoft, 1 Aug 2026
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM AgentsarXiv, 3 Jun 2026
- Hidden in Memory: Sleeper Memory Poisoning in LLM AgentsarXiv, 14 May 2026
- MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM AgentsarXiv, 16 Jul 2026
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025), section 3.5NIST, 1 Mar 2025