Definition · AI security
AI red teaming
AI red teaming is the practice of running authorised, adversary-emulating exercises against an AI system to find weaknesses before an attacker does. AI red teaming tests the deployed system as a whole, including models, data, agents, tools, retrieval and the people around them, rather than scoring a model against a fixed benchmark.
Last reviewed
Key points
- MITRE ATLAS defines AI red teaming as recurring, authorized and threat-informed exercises that simulate realistic adversary behaviour, run before deployment and throughout operation rather than once.
- The scope is the system, not the model. ATLAS says exercises should cover models and data, agents with their memory and tools, retrieval, identities and permissions, dependencies, infrastructure and human workflows.
- An ATLAS mapping is a test instruction, not advice. Against training data poisoning it says to introduce controlled poisoned records into representative pipelines, then verify provenance, sanitization, review, drift detection, model validation and rollback.
- Microsoft's AI red team emulates benign users as well as attackers, because an AI system can fail badly with nobody attacking it. Classical red teaming assumes an adversary.
- One term covers two jobs. NIST's generative AI profile counts members of the public probing for harmful or discriminatory outputs; ATLAS means threat-informed adversary emulation by a security team.
MITRE ATLAS files AI red teaming as mitigation AML.M0035, and in the 2026.08 collection it maps to 33 techniques — more than any other mitigation there, the next highest being 23. Each mapping is a test instruction, not advice. Against data poisoning ATLAS says to “introduce controlled poisoned records or triggers into representative data pipelines” and then verify provenance, sanitization, review, drift detection, model validation and rollback. Against model extraction it says to simulate functional extraction through inference queries and establish authentication, rate limits, output restrictions and anomaly detection. Against prompt injection it says to test direct, indirect and triggered instructions through user input, retrieved data, documents, images, metadata and tool output. The exercise runs in three phases — plan and scope, execute, then assess and remediate — and repeats whenever the system or the threat landscape changes.
How it differs from classical red teaming
The attacker needs less. Microsoft, reporting on over 100 generative AI products it has red teamed, found that “basic” techniques often work as well as, and sometimes better than, gradient-based methods. Those methods “typically require full access to the model, which most commercial AI systems do not provide”.
The attacker may not exist. Microsoft’s threat model “does not assume adversarial intent” and emulates benign users who trigger failures by accident — a category classical adversary emulation has no slot for.
And the finding is slipperier. Microsoft observes that traditional vulnerabilities are “usually reproducible, explainable, and straightforward to assess in terms of severity”. A single harmful completion is none of those.
Where definitions disagree
ATLAS means a security function. NIST’s generative AI profile means something broader, listing AI red-teaming as a form of structured public feedback beside focus groups and field testing, looking for “inaccurate, harmful, or discriminatory outputs”, and naming a “General Public” type performed by users “not necessarily AI or technical experts”. OWASP’s guide spans four areas: model evaluation, implementation testing, infrastructure assessment and runtime behaviour analysis. Feffer and colleagues surveyed the practice and found it diverging on purpose, artifact, setting and outcome, warning that public appeals to red-teaming “as a panacea for every possible risk verge on security theater”. Ask which of these a vendor means before reading their report.
Questions and answers
What is AI red teaming?
AI red teaming is the practice of attacking your own AI system, under authorisation, to find out how it fails before someone else does. MITRE ATLAS describes it as recurring, authorized and threat-informed exercises that simulate realistic adversary behaviour, run before deployment and repeated throughout operation. The target is the whole deployed system — models and data, agents with their memory and tools, retrieval, permissions, dependencies and the human workflows around them — not the model alone.
How is AI red teaming different from classical red teaming?
Three things change. The attacker does not need model expertise. Microsoft's red team reports that basic techniques often work as well as, and sometimes better than, gradient-based methods. The actor is not always hostile, because Microsoft emulates benign users who trip system failures unintentionally, which classical adversary emulation does not model. And the findings are harder to pin down. Microsoft notes that traditional security vulnerabilities are "usually reproducible, explainable, and straightforward to assess in terms of severity", where a prompt that produced harmful output once may not explain itself or repeat.
Is AI red teaming the same as benchmarking a model for safety?
No, and Microsoft makes this one of its eight lessons. Safety benchmarks score a model against datasets that encode harms already known, so they measure "preexisting notions of harm". A red team goes after the unfamiliar cases, and Microsoft describes its own team helping to define novel harm categories and build new probes for them. Benchmarks tell you how a model scores on last year's questions; red teaming tells you how this year's system breaks.
Who should do AI red teaming?
It depends on which definition you are working to. NIST's generative AI profile lists four types — General Public, Expert, Combination, and Human / AI — and says the quality of results is related to the background and expertise of the team, noting that demographically and interdisciplinarily diverse teams "can be used to identify flaws in the varying contexts where GAI will be used". ATLAS describes something narrower, a standing and threat-informed security function with rules of engagement, stop conditions and a remediation loop.
Does AI red teaming make an AI system secure?
It does not, and Microsoft's eighth lesson says so directly — "the work of securing AI systems will never be complete". The stated goal is to raise the cost of attacking the system, using break-fix cycles that repeat rounds of red teaming and mitigation, an approach Microsoft applied to safety-align its Phi-3 models. ATLAS agrees on the shape. Red-teaming is continuous and should be repeated whenever the threat landscape or the system changes.
Sources
- NIST Computer Security Resource Center glossary, Red Team (CNSSI 4009-2022)NIST
- MITRE ATLAS, AML.M0035 AI Red Team (collection 2026.08)MITRE
- Lessons From Red Teaming 100 Generative AI ProductsMicrosoft, via arXiv, 13 Jan 2025
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1, July 2024), Appendix ANIST
- Red-Teaming for Generative AI: Silver Bullet or Security Theater?arXiv, 29 Jan 2024
- GenAI Red Teaming GuideOWASP, 22 Jan 2025