Definition · AI agents
Oracle poisoning
Oracle poisoning is an attack that corrupts a structured knowledge graph an AI agent queries at runtime through a tool-use protocol, so the agent reaches incorrect conclusions through otherwise correct reasoning. It targets the data an agent looks up while acting, not the instructions it is given.
Last reviewed
Key points
- Oracle poisoning corrupts data, not instructions. The paper that coined the term states it directly: unlike prompt injection, it manipulates the data agents reason over.
- It targets a live, structured store an agent queries through a tool-use protocol at runtime, not a document a retrieval system embeds and not the agent's own persisted memory.
- The result is described as incorrect conclusions through correct reasoning. The model's logic is not broken; the facts it reasoned from were.
- In the paper's core test, 269 of 270 trials across nine models had the model accept a fabricated security claim under a directed query.
- Of five defences tested, only read-only access control closed the direct mutation route; the other four were partial and varied by model.
How it works
An AI agent that uses tools does not just read a prompt. Mid-task, it queries a structured store — a code knowledge graph, a dependency index, a permissions database — through a tool-use protocol, and treats what comes back as fact. Oracle poisoning corrupts that store before the query happens. The paper that coined the term formalises this as modifying a knowledge graph so that an agent’s conclusions change between the clean and the poisoned version, while its reasoning over the poisoned version stays internally consistent. Nothing about the model’s logic breaks. The facts underneath it do.
The paper tested this against a production 42-million-node code knowledge graph, using six scenarios ranging from squatting a package version to quietly editing properties on nodes that already exist. Under directed queries that pointed nine models straight at the poisoned data, 269 of 270 trials accepted a fabricated security claim; under open-ended prompts, where the model had to decide on its own whether to trust the graph, acceptance ranged from 3 to 55 percent.
Why it matters
An agent that queries a poisoned graph is not malfunctioning by any test that watches its instructions or its output format. It followed its instructions, called its tools correctly, and reasoned soundly from what those tools returned. Prompt-injection defences look for text that reads like a command; a poisoned graph entry does not have to read like anything unusual; it just has to be false, and the agent had no way to check.
Of five defences the paper tested, only read-only access control closed the direct route — an attacker who cannot write to the graph cannot poison it this way. The other four were partial and varied by model.
In practice
The paper’s six scenarios show how varied the entry points are: squatting a package’s version history, injecting a fake link in a dependency chain, fabricating a sanitiser function in the call graph so a vulnerable pattern looks safe, rerouting telemetry, seeding a poisoned pattern for an AI code generator to copy, and — in the sparest case — changing properties on nodes that already exist, adding nothing new for a diff to catch.
The four defences that were not a full fix each failed a different way: graph history preservation missed property modification, because nothing new was added to flag; audit logging caught the tampering only after it had already been acted on; a devil’s-advocate prompt that told the model to actively look for fabrications caught them at a rate that swung from 43.3 to 100 percent depending on which model was running it; and hardening the system prompt against the attack had no measurable effect at all. A survey of four other code-intelligence platforms suggests the exposure is not one vendor’s bug: any agent that trusts a structured store without checking where an entry came from is a candidate.
Where definitions disagree
The term is new. The paper that coined it, published by Microsoft Research researchers in May 2026, is — as of this review — the only source using “oracle poisoning” for this attack class; no independent glossary or vendor academy has an entry for it yet, so there is no competing definition to reconcile.
The paper does take care to place the term against two neighbours that are easy to confuse it with. Against RAG poisoning, it argues oracle poisoning bypasses any embedding layer and operates on raw structured query results rather than corrupting the documents a retrieval system embeds. Against agent memory poisoning, it frames the difference as target: memory poisoning delivers persistent content into an agent’s own store, while oracle poisoning corrupts an external structured store the agent merely queries. Whether the wider field adopts “oracle poisoning” as the name for this class, or folds it into a broader poisoning taxonomy, is not yet settled.
Questions and answers
What is oracle poisoning?
Oracle poisoning is an attack that corrupts a structured knowledge graph an AI agent queries at runtime, so the agent's own reasoning stays consistent but reaches a false conclusion because the facts it looked up were tampered with.
How is oracle poisoning different from prompt injection?
Prompt injection hides instructions in text a model reads, hijacking what the model does. Oracle poisoning tampers with the data a model looks up while acting, so nothing tells it to do anything different; it just reasons correctly from false facts.
Does read-only access control stop oracle poisoning?
It stops the direct route, where the attacker writes to the graph themselves. The paper that coined the term tested it against four other defences, and it was the only one that closed that vector entirely; the others were partial or depended on which model was doing the reasoning.
Sources
- Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent ReasoningarXiv, 1 May 2026
- Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent ReasoningarXiv, 1 May 2026
- Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning (publication listing)Microsoft Research, 1 May 2026