What matters in AI.

Subscribe

Learn / AI security

Definition · AI security

Indirect prompt injection

Indirect prompt injection is an attack in which instructions are hidden in third-party content — a web page, an email, a calendar invite, a log line, a PDF or an issue comment — that a language model later reads. The user never supplies or sees the instruction, and the model carries it out with the operator's privileges.

Last reviewed

Key points

  • OWASP holds indirect prompt injection inside LLM01, the top prompt-injection risk in both its 2025 and 2026 LLM Top 10 lists, as the case where content the model ingests acts as the injection.
  • It works because a model treats instructions and data as one token stream: content the model is asked to process can also instruct it, and nothing marks the difference.
  • The attacker never has to breach the application. They place text where the model will read it, and the model, operating with the developer's privileges, does the work.
  • In 2026 it crossed from proof of concept to live exploitation: Google measured a 32 percent relative rise in malicious IPI pages, and documented payloads forced payments, file deletion and API-key theft.
  • Defence is architectural, not interceptive: least privilege, human approval for risky actions, and isolating untrusted content bound what a bypassed model can do.

Indirect prompt injection is prompt injection delivered through content the model reads rather than words the user types: OWASP defines it as the model ingesting content that acts as prompt injection, where the user did not supply or see those instructions.

How it works

The attacker plants instructions in third-party content — a web page, email, calendar invite, log line, PDF, issue comment. When the model ingests that content in a normal task, it reads the instruction as a command.

It works because a model keeps instructions and data on one token stream; OWASP states there is no architectural distinction between the two. The user is the weak spot: the instruction is never seen, so a human-in-the-loop review cannot veto it.

The attacker never breaches the application. OWASP: the attacker does not need to compromise the backend directly; they place text where the developer’s model will read it, and the model does the work with the developer’s privileges. Microsoft notes the AI’s inability to distinguish user input from external content is what makes input validation insufficient.

Why it matters

In 2026 it crossed from proof of concept to live exploitation; CSA cites Google and Forcepoint telemetry, with Google measuring a 32 percent relative rise in malicious IPI pages between November 2025 and February 2026. Documented payloads forced a $5,000 PayPal transfer, recursive file deletion by coding agents and API-key theft.

The consequence is an action, not an answer. When the model can call tools it acts with the operator’s credentials — agent hijacking — and the damage is set by the lethal trifecta and excessive agency.

No reliable prevention exists, so defence is architectural: least privilege, human approval for risky actions, and isolating untrusted content bound what a bypassed model can do.

In practice

ThreatLabz tested the payment-scam pattern across 26 models: four executed the payment. The attack hid instructions in fake API documentation using JSON-LD structured data and text positioned off-screen, invisible to a human but present in the DOM. In the GrafanaGhost case, an attacker poisoned Grafana log entries so the AI assistant treated embedded text as instructions when summarising logs, and the data left through Markdown image rendering. Neither that log injection nor most AI-specific flaws received a CVE, so version pinning and vendor changelogs carry weight where scanner signatures cannot.

In hiring, the same attack appears as résumé prompt injection: text an applicant hides in a résumé so an AI screener reads it and a human reviewer does not.

Questions and answers

What is the difference between direct and indirect prompt injection?

In direct injection, a user or attacker types instructions into the model. In indirect injection, the attacker plants instructions in third-party content — a web page, email, calendar invite, log line, PDF or issue comment — that the model reads later as part of a normal task. The user never supplies or sees the instruction. OWASP ranks the indirect form within LLM01 and calls it the more dangerous case, because the attacker never has to touch the application.

Can indirect prompt injection be prevented by better prompting?

No. OWASP states no reliable prevention mechanism exists today, because a model makes no architectural distinction between instructions and data. Instructing the model to ignore injected instructions is text in the same channel the attack uses. Defence is architectural: limit what the model can do, require human approval for risky actions, and isolate untrusted content so a bypassed model has no reach.

Has indirect prompt injection happened outside research?

Yes. CSA reported in April 2026 that it crossed from proof of concept to live exploitation, citing Google and Forcepoint telemetry on pages seeded with hidden instructions. Documented payloads forced a $5,000 PayPal transfer, recursive file deletion by coding agents and API-key theft. Zscaler ThreatLabz showed a 2026 payment scam that induced four of 26 tested models to send funds.

Is indirect prompt injection the same as agent hijacking?

Agent hijacking is indirect prompt injection aimed at a system that can call tools, so the result is an unauthorised action taken in the operator's name rather than an unwanted answer. The injection mechanism is the same; what changes is the blast radius, which OWASP ties to whether the model's output drives tool calls.

Sources

  1. LLM01:2026 Prompt InjectionOWASP GenAI Security Project, 4 Aug 2026
  2. Defend against indirect prompt injection attacksMicrosoft Learn, 19 Mar 2026
  3. Indirect Prompt Injection Goes OperationalCloud Security Alliance AI Safety Initiative, 26 Apr 2026
  4. Indirect Prompt Injection in Web Content Targets AI AgentsZscaler ThreatLabz, 2 Jul 2026

Guides that use this term