What matters in AI.

Subscribe

Learn / AI security

Guide · AI security

How to red team an LLM application

To red team an LLM application, start from the damage the application could do in its real setting and work back to the attacks that would cause it. Test the deployed application, not only the model, including every place untrusted text enters. Probe by hand first, then automate. Run each attack more than once, and test again after every fix.

Last reviewed

This guide applies AI red teaming to one application built on a language model.

Start from the damage, not from a jailbreak list

Pick the outcomes you cannot accept before you pick a single attack. Microsoft’s AI red team, reporting on more than 100 generative AI products, says that “Starting from potential downstream impacts, rather than attack strategies, makes it more likely that an operation will produce useful findings tied to real world risks”. Work backwards from each outcome to the paths an attacker could take.

Two questions set those outcomes: what the application can do, and where it is used. Microsoft’s example is the same model used either as a creative writing assistant or to summarise patient records. The second “clearly poses much greater downstream risk”. An assistant that can send email, run code or read other users’ data has more to lose than one that only chats.

Then write the scope down. OWASP’s GenAI Red Teaming Guide says it should name “which models and systems will be tested, what types of tests will be conducted, and what areas or activities are explicitly excluded”. Its rules of engagement include where testing runs, who to escalate to, how test data is handled, how to roll back, and who has given permission.

Attack the application, not only the model

A model that holds up on its own can still sit inside an application that does not. Microsoft warns that “red teaming strategies that target only models may not translate into vulnerabilities in production systems”, and that strategies ignoring the components that are not generative AI, such as “input filters, databases, and other cloud resources”, “will likely miss important vulnerabilities”. In one of its case studies the flaw was an outdated FFmpeg library in a video-processing app. A crafted video file could make the server reach internal resources, a server-side request forgery. No prompt was involved.

OWASP splits the work into layers you can test one at a time:

  • the model: alignment, robustness and bias, plus where the model came from: provenance, malware in the model, and the training data pipelines
  • the implementation: getting past supporting guardrails, such as those “included in a system prompt”, poisoning the data a retrieval augmented generation pipeline retrieves, and testing controls such as model firewalls and proxies
  • the system: the components around the model, excessive agency and the supply chain
  • runtime: business processes, over-reliance and social engineering

Test where users are. Microsoft’s planning guide says to test “on the production UI as much as possible because this most closely resembles real-world usage”, and to say in the report which endpoint each result came from. A result from the raw model API does not tell you what your AI guardrails and system prompt do.

Go where untrusted text gets in

Content the application reads but the user did not type is attack surface too. Indirect prompt injection puts the attack there. Microsoft’s examples are documents in a retrieval pipeline and an email that a copilot summarises, and it says hiding instructions in documents has let it “alter model behavior and exfiltrate private data” in “a variety of operations”. Plant test instructions in each channel the application reads, not only in the chat box.

Then chain the steps. One Microsoft operation used prompt injections written in a low-resource language, one with little text available for training, to discover the app’s internal Python functions. A second injection, carried in content the app read, got the model to write a script calling those functions. Running that script pulled out private user data. Microsoft notes that these injections “were crafted by hand and relied on a system-level perspective”.

OWASP’s prompt injection entry gives the stance to test from: treat “the model as an untrusted user” and check whether trust boundaries and access controls hold when it misbehaves.

Probe by hand, then automate what you learned

Microsoft’s planning guide calls for “an initial round of manual red teaming before conducting systematic measurements”. Start open-ended, with testers free to report anything that goes wrong. Turn what they find into a list of harms, then run guided rounds against the list and add new harms as they appear. Mix testers who think like attackers with ordinary users who had no part in building the app, and give security specialists the jailbreaking and system prompt extraction work.

Automation comes after. Anthropic describes experts first probing by hand, then standardising their attacks, and then using a language model “to generate hundreds or thousands of variations of those inputs”. Open-source frameworks exist for this. Microsoft’s PyRIT supplies prompt datasets, prompt converters such as encodings, automated attack strategies and scorers. garak, from Derczynski and colleagues, runs probes that are each “designed to elicit a single kind of LLM vulnerability”.

Keep a person in the loop. Microsoft says these tools “should not be used with the intention of taking the human out of the loop”. Using a model to score outputs works for simple checks, but it is “less reliable” in medicine, cybersecurity and chemical, biological, radiological and nuclear (CBRN) weapons, where only subject matter experts can judge the output. garak’s authors say it is “designed to be used as part of human assessment”, and that it “does not deal with security issues presenting in a broader system context, such as code execution or insufficient access controls”.

Run every attack more than once

The same prompt can fail and then succeed. OWASP says “a prompt that fails initially may succeed upon repeated attempts”, so it recommends several attempts per attack and a success threshold based on repeated trials. For prompt injection OWASP sets a lower bar: “a single successful adversarial response may indicate a vulnerability”, though it is worth confirming it reproduces. Microsoft makes the same point about scale: running attacks in volume helps “estimate how likely a particular failure is to occur”.

Record enough to rerun each finding. Microsoft’s guide lists the date, a unique ID for the input and output pair where one is available, the prompt, and a description or screenshot of the output.

What the report can and cannot claim

One finding shows that a failure can happen. On its own it does not show how common the failure is. Microsoft’s planning guide says red teaming is “not a replacement for systematic measurement”, and warns: “It is important that people do not interpret specific examples as a metric for the pervasiveness of that harm.” Repeating an attack gives a rough rate for that attack. The guide treats red teaming as the step that produces the list of harms to measure and mitigate next. garak’s authors make a related point about their own tool: “It would be pointless to attempt to treat garak results as a benchmark”, because the framework “is customizable in each run” and its output “would (and should) vary for different contexts”.

Retest after every fix. Microsoft’s planning guide says to test with and without the mitigations in place, and the AI red teaming topic covers why the loop never ends. A clean run proves only that these testers found nothing with these attacks on this version.

Sources

  1. Lessons From Red Teaming 100 Generative AI ProductsMicrosoft, via arXiv, 13 Jan 2025
  2. Planning red teaming for large language models (LLMs) and their applicationsMicrosoft Learn, 13 May 2026
  3. GenAI Red Teaming Guide, Version 1.0OWASP GenAI Security Project, 23 Jan 2025
  4. LLM01:2025 Prompt InjectionOWASP GenAI Security Project
  5. garak: A Framework for Security Probing Large Language ModelsarXiv, 16 Jun 2024
  6. Challenges in red teaming AI systemsAnthropic, 12 Jun 2024

Terms in this guide