Definition · AI agents
ObliInjection
ObliInjection is a prompt-injection attack for LLM agents whose input is assembled from multiple sources, where the attacker controls only some segments and does not know their final order. It optimizes an "order-oblivious loss" so a contaminated segment succeeds regardless of where it lands in the assembled input.
Last reviewed
Key points
- ObliInjection targets LLM applications whose input is assembled from multiple sources, where the attacker controls only some segments and does not know their final position in the assembled input.
- Two named contributions carry the coinage — the order-oblivious loss, which scores a candidate attack by its expected success averaged over random segment orderings, and orderGCG, an optimization algorithm built to minimize that loss.
- Averaged across seven LLMs, it reached attack success rates of 99.0%, 98.7% and 99.6% on three datasets; on Amazon Reviews the strongest prior attack reached 56.8%.
- It works with as little as one of 6 to 100 segments contaminated.
- Fine-tuning defenses that separate instructions from data (StruQ, SecAlign) cut its success to 78% and 64% on Llama-3-8B for Amazon Reviews; leave-one-segment-out, delimiters and both detection methods tested left it largely intact.
How it works
Many LLM applications assemble their input from several sources at once — retrieved documents, tool outputs, reviews — and concatenate the segments before the model sees them. An attacker who controls only one or a few segments has no way to know where the pipeline will place them. Earlier prompt-injection attacks assumed the attacker controls the whole input, or optimized for one fixed position and lost effectiveness elsewhere.
ObliInjection optimizes directly against that uncertainty. Its order-oblivious loss scores a candidate segment by its expected success averaged over random orderings of the clean and contaminated segments, not by its success at one position. orderGCG is the algorithm built to minimize that loss: it extends the token-substitution search of standard GCG by accumulating approximate loss estimates for each candidate across iterations, offsetting the extra variance from sampling different shadow segments and orderings on every step.
Why it matters
A multi-source agent call is the ordinary shape of a production RAG pipeline or tool-using agent, not an edge case. Averaged across seven LLMs, ObliInjection reached attack success rates of 99.0%, 98.7% and 99.6% on three datasets while contaminating as little as one segment out of 6 to 100. On Amazon Reviews the strongest prior attack reached 56.8%.
Some defenses barely helped: removing one segment at a time and adding delimiters between segments left success at 99.3% and 96.3%. Two detection methods missed most contaminated segments — a perplexity-based detector had a 92.6% false-negative rate, DataSentinel had 79.6% — and the contaminated segments still succeeded 85.5% to 100% of the time across five LLMs.
In practice
Fine-tuning defenses did better, but not completely. Tested on Llama-3-8B for Amazon Reviews, StruQ, which trains the model to distinguish instructions from data, cut success to 77.9%, and SecAlign to 63.8%. Fine-tuning SecAlign further on ObliInjection’s own samples still left 53.3% when those samples matched the attack setting, and 81.0% against a slightly different attack configuration. None of the defenses tested eliminated the attack.
Where definitions disagree
The term is new: the paper that coins it, by Reachal Wang, Yuqi Jia and Neil Zhenqiang Gong of Duke University, was accepted to NDSS Symposium 2026 from a December 2025 arXiv submission, and is the only source using “ObliInjection” for this technique so far. It is not a wholly new category: prompt injection and indirect prompt injection name the general mechanism and the delivery channel, and ObliInjection is a specific optimization technique within that mechanism, built for the case where several legitimate sources are concatenated in an order the attacker cannot predict. The paper separates it from some RAG poisoning attacks, such as PoisonedRAG, by the number of contaminated segments: those assume the attacker’s texts are a majority of the retrieved passages, while ObliInjection needs only one.
Questions and answers
What is ObliInjection?
ObliInjection is a prompt-injection attack for LLM agents whose input is assembled from multiple sources. It works even though the attacker controls only some of the segments and does not know where in the final input their segment will land, by optimizing against an "order-oblivious loss" that scores success averaged over possible segment orderings.
How is ObliInjection different from ordinary prompt injection?
Ordinary prompt injection and indirect prompt injection assume the attacker's text is either the whole input or a single identifiable segment. ObliInjection is built for the case where several sources are concatenated into one context — several retrieved documents or tool outputs, say — and the attacker's contaminated segment could land anywhere in that concatenation, in an order the attacker cannot predict.
How much of the input does an attacker need to control for ObliInjection to work?
The paper that coined the term reports it works with as little as one of 6 to 100 segments contaminated, depending on the dataset, reaching average attack success rates above 98% across the three datasets tested.
What stops ObliInjection?
Fine-tuning defenses fared better in the paper's tests — StruQ and SecAlign cut its success rate to 78% and 64% on one model and dataset — while removing one segment at a time, adding delimiters between segments, and the two detection methods tested left it largely unaffected.
Sources
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source DataarXiv, 10 Dec 2025
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source DataarXiv, 10 Dec 2025
- ObliInjection paper pageNDSS Symposium, 1 Feb 2026