What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Agent data injection

Agent data injection is a class of indirect prompt injection in which an attacker disguises malicious data as trusted metadata, such as element identifiers, comment authorship or tool call records. The agent still performs the user's task, but acts on attacker-controlled data, causing unintended clicks, command execution or merged code.

Last reviewed

Key points

  • Agent data injection (ADI) is indirect prompt injection that forges trusted data rather than hidden instructions. The agent keeps doing the user's task, but on attacker-controlled identifiers, author names or tool results.
  • It works by probabilistic delimiter injection — characters a tool treats as text but a model reads as structure — so untrusted content is absorbed as if it were trusted metadata.
  • Six frontier models were tested and all were vulnerable. Attack success ranged from 31.3 to 43.3 percent on JSON data and 33.3 to 100 percent on web page content.
  • It defeats the standard prompt-injection guardrails. Input guardrails gave no improvement, while instruction injection against the same defenses succeeded under 1 percent of the time.
  • The root cause is architectural. Agent frameworks do not separate trusted from untrusted data inside the agent context, only instructions from data.

How it works

An agent does not see raw data, it sees a rendering: a web agent turns a page into elements with identifiers for “click that one”; a coding agent turns issues and pull requests into a record of who wrote what and which tools ran. ADI attacks that rendering, not the instructions.

The technique is probabilistic delimiter injection. Agent data is separated by delimiters — braces and quotes in JSON, tags around a tool call and its result — which a parser matches exactly, but a model infers from what they look like, accepting delimiters a parser would reject.

An attacker who controls one untrusted field — a product review, an issue comment — injects characters imitating those delimiters. The model concludes a new field exists where the tool put only text, and whatever it claims to be — an element identifier, a maintainer’s name, a passed test — is absorbed as trusted data.

The agent still does what the user asked — clicks a button, applies a fix, merges a pull request — just the wrong button, someone else’s fix, or tests that never ran.

Why it matters

Instruction injection needs the model to abandon the user’s task. ADI does not, so existing defenses do not carry over: they look for untrusted content behaving like an instruction. Guardrails and plan-then-execute interrupt that one pattern. ADI uses no imperative language, so a forged sender name reads as ordinary data with nothing to flag — in the 2026 study, guardrails gave no measurable improvement, while the same mitigations cut instruction injection to near zero.

The consequence scales with the agent’s reach: a misclicked “Buy Now” is a purchase, a spoofed maintainer comment is remote code execution, a fabricated test result an unreviewed commit. The attacker never breaches the agent; they edit content it was already going to read.

In practice

The published study demonstrated the attack against deployed agents from Anthropic, OpenAI and Google.

  • Web agents. A crafted product review reused a sequential, predictable element identifier, so the agent clicked the real “Buy Now” button while trying to read more reviews. Claude in Chrome, Antigravity and Nanobrowser were vulnerable; Atlas resisted because its identifiers are random per page load.
  • Coding agents. A GitHub comment spoofed its author as a maintainer, so an agent asked to “apply the maintainer’s fix” ran an attacker’s command instead. Claude Code, Codex and Gemini CLI were affected.
  • Pull request review. A fabricated tool-result in a PR description made an agent believe a malicious commit had already been reviewed, and merge it.

Trade-offs

Two controls cut the attack sharply, and both cost something.

Runtime-randomized identifiers, as in Atlas, cut success from 49.1 to 28.7 percent, but only help where an identifier exists to randomize — not for unstructured content like a comment body.

Strict data-flow tracking, keeping trusted and untrusted data in labelled channels instead of letting the model infer the boundary, eliminated the attack but halved task completion, to 36.5 percent against an 86.5 percent baseline.

Sandboxing and dual-LLM designs helped but did not close the gap. The fully effective defense is too expensive for many deployments today, and the partial ones reduce rather than remove the risk — treating prompt-injection coverage as sufficient for agent risk is the mistake this research argues against.

Questions and answers

How is agent data injection different from prompt injection?

Prompt injection smuggles an instruction into untrusted content, so the agent performs the attacker's task. Agent data injection forges trusted metadata, so the agent performs the user's task with attacker-controlled data. The agent's task stays the same, which is why instruction-focused defenses miss it.

Do prompt-injection guardrails stop agent data injection?

Largely no. In the 2026 study, input guardrails produced no improvement over no defense at all, and ADI succeeded at up to 50 percent against a suite of mitigations that reduced classic instruction injection to under 1 percent.

What actually defends against agent data injection?

Strict data-flow tracking between trusted and untrusted sources eliminated the attack in testing, and runtime-randomized element identifiers cut success roughly in half. Both are architectural controls rather than prompt-level ones, and the fully effective option carried a large utility cost.

Sources

  1. Agent Data Injection Attacks are Realistic Threats to AI Agents (arXiv:2607.05120v1)arXiv, 6 Jul 2026
  2. Agent Data Injection: A New Attack Class Beyond Prompt InjectionCloud Security Alliance, 17 Jul 2026

Guides that use this term