What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Memory provenance

Memory provenance is the practice of labelling every entry an AI agent writes to persistent memory with where that content came from, and keeping enough history to quarantine or roll back a suspect entry. MITRE ATLAS names it as one control within its memory hardening mitigation, AML.M0031.

Last reviewed

Key points

  • MITRE ATLAS names memory provenance as one of four memory hardening controls, AML.M0031. The control is to record the source of all memory updates, preserve a history with known good versions, and quarantine or roll back suspicious records.
  • The label goes on at write time. Microsoft's Zero Trust guidance asks for source, identity, timestamp and model version on every memory entry, and gates the write on user intent as well as authority.
  • The record is what makes containment complete. ATLAS says memory hardening lets poisoned records be identified, quarantined and rolled back, and only the origin says what else shares its source.
  • Memory provenance is not dataset provenance. ATLAS separates them by lifecycle phase, placing dataset provenance in data preparation, memory hardening in model engineering, deployment and monitoring. Training-time record versus runtime state.
  • No source here argues the write-time record is unnecessary; the argument is over what inspection adds. Microsoft inspects content at the write and the read; ATLAS scopes content filtering out of memory hardening.

An agent’s memory is written out of ordinary content: a tool result, a fetched page, an email body. A later session reads it back as the agent’s own prior note. Memory provenance keeps the origin attached, so what the agent was told stays distinct from what it worked out.

What the record carries

Microsoft’s Zero Trust guidance names the fields: “Label provenance on every memory entry: source, identity, timestamp, model version.” It also gates the write, asking teams to “avoid implicit or autonomous memory creation from untrusted sources”.

MITRE ATLAS adds what makes the label useful later. Its control reads “Record the source of all memory updates and preserve a history with known good versions. Quarantine or roll-back suspicious records.” A label with no history behind it marks the entry but leaves nothing to restore.

A 2026 survey of agent provenance adds the step between. Memory writes “often transform information: documents are summarized, observations are merged, preferences are inferred”, so the record carries the transformation too, distinguishing “raw observations, extracted facts, inferred memories, and revised memories”.

Why it matters

Without the origin, an entry is only its text. The survey puts it plainly: if a memory is stored without source and transformation metadata, “later verification can inspect only its surface content, not its evidential basis”. Surface content is the part an attacker writes.

The label also makes containment complete. ATLAS states both halves in its agent memory poisoning mapping: hardening prevents “saved data from becoming higher-authority instructions” and enables “poisoned records to be identified, quarantined, and rolled back”. The first is blanket, like Microsoft’s “candidate context, not authoritative truth”, and needs no label. For the second, monitoring can flag one entry; the label says what else shares its source.

In practice

The tooling is early. ATLAS names OWASP Agent Memory Guard and the Microsoft Agent Governance Toolkit as open-source implementations of memory hardening controls; the OWASP one is an Incubator project whose own roadmap puts applying for Lab promotion at the end of 2026. The survey is blunter about the research systems. Across the memory systems and security frameworks it tabulates, “write and retrieval support are common” while “provenance-specific capabilities such as source attribution, conflict handling, staleness detection, contamination tracking, and evidence-aware verification remain limited”. Source attribution is the subject of this page, and the survey puts it in that list.

One architecture puts the ordering first. Eywa, a May 2026 preprint, describes a memory built “around evidence before belief”: immutable source evidence is stored before any canonical fact is derived from it, extracted memories are validated against source support, and retrieved context is “returned separately from answer instructions”. Its published numbers are answer-quality benchmarks, not a measurement of how well it resists poisoning.

One practitioner goes further than the standards do. Writing about persistent memory on Hacker News in June 2026, they described treating “provenance as first-class in the stored state, not a tag I hope survives”, with every stored line carrying where it came from and a read rule that “outside-origin content is quotable as fact but never executable as instruction”. They add a third rule that no standard checked here states: “never summarize across the trust boundary”, so a foreign sentence “gets stored verbatim and tagged, or it does not get stored”. ATLAS does govern the result: M0031 names “conversation summaries” among the persistent state its four controls protect. Of the published guidance Microsoft comes nearest to the operation, and only in an example: illustrating what a data handling taxonomy might contain, it offers “Never infer”, which says of sensitive attributes to “only add to memory if explicitly provided by the user” — so the agent may store what it was told and not what it worked out. Even that line is drawn around a class of content rather than around a trust boundary, and neither it nor ATLAS reaches summarising. Treat the practitioner’s rule as field practice, not as guidance.

The write path is the ordinary one. ATLAS notes that agent memory is controlled through normal conversation, so an adversary can inject memories by direct or indirect prompt injection. Nothing is exploited at the write. That is what a provenance record is for: it does not stop the entry being made, it keeps the entry answerable afterwards.

The assumption being protected is the same one instruction privilege escalation turns into an attack in a different store. That paper finds that the two defences agent harnesses rely on, an agent trained to refuse unsafe instructions and a separate permission reviewer, both rest on “a crucial assumption: the role labels faithfully reflect the true provenance of the content”, and that a harness “may reconstruct context without preserving the content’s original provenance”. In memory the loss happens once, at the write, and every later session inherits it.

Where definitions disagree

Whether the control includes content inspection. ATLAS and Microsoft both ask for the label at the write. They part over whether judging the content is part of the same control. ATLAS scopes it out of memory hardening and says so directly: “Unlike guardrails that evaluate content during an interaction, memory hardening controls how durable agent state is created, modified, isolated, audited, and recovered.” Filtering for malicious prompts is routed to a separate mitigation, Generative AI Guardrails, which ATLAS says does not replace provenance. Microsoft puts content inspection at both ends of the memory path. Its write gate is “also a good time to sanitize inputs using a data handling taxonomy”, and its guidance table puts that more firmly, recommending teams “Classify or govern data at write time; block inappropriate or harmful data before writing to memory”; at retrieval it asks teams to “Revaluate for sensitive or malicious content (for example, Prompt Shields)” and to apply “retrieval-time Prompt Shields to detect indirect attacks before injecting memory into reasoning context”.

Part of Microsoft’s answer, then, is that some of it never gets written. The Hacker News practitioner reports flatly that the read-time half does not stand alone: “a single read-time filter doesn’t either, because by next session the foreign sentence no longer looks foreign”. The survey supplies a research version of the same claim, reporting that NeuroTaint “emphasizes that unsafe influence can survive semantic transformation and memory reuse rather than appearing as exact text reuse”. Nobody in this set argues the write-time record is unnecessary; the argument is over what inspection adds, and at which end of the path it belongs.

How much the term covers. ATLAS uses memory provenance narrowly, as one bullet in a four-control mitigation whose formal name is Memory Hardening. The survey uses it as the name of a property of the whole memory system, covering writes, retrieval traces, temporal validity, conflict status and downstream influence, and titles a section “Memory as Provenance-Bearing Evidence”. Eywa, a memory architecture built on the property, uses it the wider way too, covering storage, validation and the read path. Neither sense covers AI dataset provenance, which ATLAS files as a separate mitigation, AML.M0025, at the data-preparation phase.

Questions and answers

Is memory provenance the same thing as data provenance?

No, and MITRE ATLAS draws the line with its own fields rather than in prose. Maintain AI Dataset Provenance, AML.M0025, carries the lifecycle phases Business and Data Understanding and Data Preparation, and reaches every technique it maps to through the dataset, never through agent memory. Memory Hardening, AML.M0031, carries AI Model Engineering, Deployment, and Monitoring and Maintenance. Dataset provenance is a record of what a model trained on. Memory provenance is a record of what a running agent wrote down about you last week.

What has to be recorded on a memory entry?

Microsoft's Zero Trust guidance names four fields to label on every entry, being source, identity, timestamp and model version, and asks that every create, read, update and delete be logged with identity, timestamp, source and provenance. ATLAS asks additionally for a history with known good versions, so a suspect record can be rolled back rather than only recognised. A 2026 survey of agent provenance adds the transformation, because a memory is often a summary or an inference rather than the text that produced it.

Can a filter at read time do the same job?

No source here argues that it can. Microsoft labels provenance at the write and inspects content at both ends, recommending that data be classified or governed at write time and that harmful data be blocked before it is written, and applying retrieval-time Prompt Shields before memory enters the reasoning context. ATLAS keeps memory hardening on the write and the audit trail, and routes content inspection to a separate mitigation which it says "does not replace memory access control, isolation, provenance, integrity, or recovery". The open question is narrower than substitution, and it is what inspection can recover after the fact. A practitioner on Hacker News reports that read-time filtering alone failed for them, because by the next session the foreign sentence no longer looks foreign, and a 2026 survey reports that NeuroTaint finds unsafe influence can survive semantic transformation rather than appearing as exact text reuse.

Does memory provenance stop memory poisoning?

Not on its own, and ATLAS does not claim it does. Provenance is one of four controls under memory hardening. The other three are enforcing memory access controls, authenticating memory operations and authorising them within the right user, tenant, agent and session scope; setting strong memory security policies, which covers limits on memory size and update frequency, validating memory integrity, and retention and deletion policies; and auditing and monitoring memory operations, which covers auditing security relevant reads and writes and monitoring for unusual update frequency and repeated self-authored updates. Integrity validation and that monitoring both flag a suspect entry without reading any label on it, so the provenance record is not the only path to a poisoned record. What it adds is the scope of the cleanup, being what else came from the same source and which stored version to restore. ATLAS adds that content filtering belongs to a separate mitigation and "does not replace memory access control, isolation, provenance, integrity, or recovery", and scopes memory hardening to AI Agent Context Poisoning and its Memory sub-technique, not to the Thread sub-technique.

Sources

  1. MITRE ATLAS, mitigation AML.M0031 Memory Hardening (collection 2026.08)MITRE
  2. Manage AI memory safety in agentic systemsMicrosoft, 3 Jun 2026
  3. From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents, section 5.1arXiv, 3 Jun 2026
  4. Eywa: Provenance-Grounded Long-Term Memory for AI AgentsarXiv, 29 May 2026
  5. When Context Gets Root: Privilege Escalation in LLM HarnessesarXiv, 27 Aug 2026
  6. Comment by sarracin0 on Hacker News item 48636793Hacker News, 22 Jun 2026
  7. OWASP Agent Memory GuardOWASP Foundation