What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

AI agent context poisoning

AI agent context poisoning is an attack that manipulates what an AI agent's model treats as trusted context — its saved memory or a running conversation thread — so a change the attacker makes once keeps shaping the agent's behavior afterward. MITRE ATLAS defines it as AML.T0080, split into a memory sub-technique and a thread sub-technique.

Last reviewed

Key points

  • One technique, two stores. MITRE ATLAS splits AI Agent Context Poisoning into Memory (AML.T0080.000), a durable per-user database, and Thread (AML.T0080.001), the running conversation itself.
  • Thread poisoning can spread to other people. ATLAS gives the example of an agent active in a shared Slack channel, where one user's message can shape how the agent treats everyone else in that channel afterward.
  • Rising context windows extend the attack. ATLAS notes that as token limits grow, larger context windows let instructions planted in a thread persist longer before they age out.
  • Demonstrated on shipping agents. HiddenLayer poisoned every new OpenClaw conversation through a config-file injection triggered by a greeting; SafeBreach kept instructions alive in Gemini's context through a poisoned Calendar invite.
  • The dedicated mitigation has a gap. ATLAS scopes Memory Hardening, AML.M0031, to the technique and its memory sub-technique only — thread poisoning is left to the general-purpose AI Red Team control instead.

How it works

Memory (AML.T0080.000) writes to a durable, per-user store that survives across sessions — the case memory poisoning covers. Thread (AML.T0080.001) instead poisons the conversation already in progress: ATLAS says instructions introduced into a chat thread cause behavior changes that persist “for the remainder of the thread,” and notes a thread “may continue for an extended period over multiple sessions” without being a saved database. Both arrive through direct or indirect prompt injection.

Thread poisoning can reach people who never triggered it. ATLAS gives the example of an agent in a shared Slack channel, where one malicious message “can influence the agent’s behavior in future interactions with others.” It also notes that as context windows grow, planted instructions simply persist longer before anything pushes them out.

Why it matters

Both sub-techniques have hit shipping agents, not just demonstrations on paper. HiddenLayer got OpenClaw to run a script that appended an injection to a config file it loads into its own system prompt; ATLAS records that “the context of all new threads became poisoned,” triggered the next time a victim greeted the assistant. SafeBreach kept instructions alive in Gemini’s context through a poisoned Calendar invite, some concealed behind a “Show more” control, shaping later turns on both web and Android.

The dedicated fix has a gap. ATLAS scopes Memory Hardening, AML.M0031, to the technique and to Memory specifically — it does not mitigate Thread. The only control ATLAS maps there is AI Red Team, AML.M0035, a general adversary-emulation exercise rather than one built for shared threads.

Questions and answers

Is AI agent context poisoning the same as memory poisoning?

No, memory poisoning is one of its two forms. MITRE ATLAS splits AI Agent Context Poisoning, AML.T0080, into Memory (AML.T0080.000), a durable per-user store that survives across sessions, and Thread (AML.T0080.001), the running conversation itself, which may itself continue across sessions but is not a separate saved database.

Can context poisoning through a thread affect people other than the attacker?

Yes, when the thread is shared. ATLAS gives the example of an agent active in a Slack channel with multiple participants: a single malicious message from one user can influence how the agent behaves toward everyone else in that channel afterward, not only toward the person who sent it.

Does memory hardening stop thread poisoning?

Not on its own. ATLAS scopes its dedicated memory hardening mitigation, AML.M0031, to the general technique and its Memory sub-technique. Thread poisoning is covered instead by AML.M0035, AI Red Team, a general-purpose exercise rather than a control built for long-lived or shared conversation threads specifically.

How does context poisoning differ from training-data poisoning?

Training-data poisoning corrupts the model's weights during training by injecting malicious examples into the dataset before the model is ever trained. Context poisoning corrupts what the model sees at inference time — its saved memory or running conversation thread — and does not require control over training data or model retraining. Context poisoning also outlives a single conversation turn without necessarily modifying any permanent artifact; training-data poisoning requires retraining the entire model to change behavior permanently.

How does context poisoning differ from MCP tool poisoning?

MCP tool poisoning targets the trusted tools an agent calls on — specifically by editing an already-installed tool's description or metadata to hide malicious instructions within it. Context poisoning targets the context the agent reads — its memory store or conversation thread — and manipulates the model's own reasoning, not the external systems it calls. An MCP-poisoned tool's hidden instructions may steer an agent toward certain actions; a context-poisoned agent may misinterpret correct data or ignore correct tools entirely because its context tells it to.

Sources

  1. MITRE ATLAS, technique AML.T0080 AI Agent Context Poisoning (collection 2026.08)MITRE

Guides that use this term