Learn / Category
Learn
AI security
87 topics in this category.
Guides
9 guides
- Guardrail bypass techniques
The named ways attackers get a language model past its safety controls, grouped by the control each one mainly gets around, plus the reconnaissance that comes first.
- How to detect shadow AI
Shadow AI can be found in proxy and firewall logs, identity-provider app consents, code repositories, cloud audit logs and scans for self-hosted models. Each source misses something, so check several and give staff a safe way to report what they use.
- How to red team an LLM application
Red team an LLM application by starting from the harm that would matter, attacking the whole deployed application where untrusted text gets in, probing by hand before automating, and repeating every attack.
- Input filtering vs output validation
Input filtering screens what goes into an LLM; output validation checks what comes out before a user or another system acts on it. Each misses what the other catches, and OWASP and MITRE ATLAS list both.
- Security risks of running LLMs locally
Running a language model locally keeps prompts away from a provider, but it puts a downloaded model file, an inference server and the model's behaviour on your own machine, and it leaves prompt injection and hallucinated output in place.
- Shadow AI vs shadow IT
Shadow AI is the part of shadow IT that involves AI tools. The shadow IT playbook still applies, but AI adds three things: prompts may train the provider's models depending on the account, AI can arrive through an update to software already in use, and the output itself can be wrong.
- State of AI security, September 2026 edition
Seven pillars of AI security, ranked by where problems are actually being found rather than by where the noise is.
- Targeted vs untargeted data poisoning
Untargeted poisoning aims to make a model worse on inputs in general. Targeted poisoning changes its answer on a few inputs the attacker picked and leaves the rest alone, which is why accuracy checks usually catch the first and can miss the second.
- White-box vs black-box attacks
A white-box attacker knows the model inside out; a black-box attacker knows little or nothing about it and usually works by sending inputs and reading outputs. The labels describe what the attacker knows, not what the attacker does, and NIST describes a gray-box range in between.
Definitions, A to Z
78 definitions
- Adaptive attack
An attack designed against one specific defence by an attacker who knows how it works, and the standard test of defences against adversarial examples.
- Adversarial example
An input crafted so a machine learning model gets it wrong at deployment time.
- Adversarial machine learning
The field that studies attacks exploiting how machine learning works, and how to build models that withstand them.
- Adversarial suffix
A string of tokens, found by automated search, that is appended to a harmful request so an aligned language model answers instead of refusing; the best-known method for finding one is GCG.
- Adversarial training
A defence that trains a model on the adversarial examples an attacker would use against it, so it learns to get them right.
- AI bill of materials
A machine readable inventory of the datasets, models and components a system depends on.
- AI dataset provenance
A maintained history of where a training dataset came from and every change made to it since.
- AI guardrails
Safety controls placed between users, tools, and a generative AI model that check prompts, retrieved content, outputs, and agent actions before they take effect.
- AI output watermarking
Hiding a machine-detectable signal in what an AI model generates, so the output can later be traced to that model, and why the signal can be removed.
- AI red teaming
Authorised adversary-emulation exercises run against a deployed AI system to find its weaknesses before an attacker does.
- AI security
Protecting machine learning systems from attack, including attacks on the model itself.
- AI security posture management
A set of vendor security tools that discover, monitor, and remediate AI-specific risk — exposed model weights, over-permissioned AI service accounts, exposed training data — across an organization's AI systems.
- AI security trade-offs
The costs an organisation accepts when it hardens an AI system against attack, and the decision NIST says the organisation must then make.
- AI supply chain compromise
An attack that reaches an AI system through a part it acquired — a model, a dataset, a framework, a container image or an agent tool.
- AI supply chain rug pull
An attack where an AI component is published as legitimate, gains adoption, and is then updated to a malicious version, which may let the adversary bypass the extra scrutiny new components often receive.
- Architectural backdoor
A backdoor built into a model's structure rather than its weights or training data, and in some designs it survives retraining.
- ASCII smuggling
Hiding text inside invisible Unicode tag characters, so a human sees nothing while a language model reads the hidden message.
- Attack success rate
The share of attack attempts in a test that achieved the attacker's goal, and why it describes the test as much as the system.
- Backdoor attack
A hidden conditional rule inside a trained model, dormant until an input carries the attacker's trigger.
- Clean-label poisoning
Poisoning a training set without touching a single label, which removes the mismatch every cheap filter is looking for.
- Confused deputy
A program tricked into using its own authority for someone who should not have it, because it cannot tell which of its principals a request comes from.
- Cost harvesting
Running up the bill for a victim's paid AI service by driving it past its normal operating capacity.
- Crescendo
A multi-turn jailbreak that opens with a harmless question about a topic and escalates a step at a time, each new prompt building on what the model itself just wrote, until the model produces content it would have refused if asked directly.
- Data poisoning
Controlling part of the data a model trains on, so the model learns what the attacker chose.
- Deepfake
Synthetic media made with generative AI to pass as authentic, which MITRE ATLAS catalogues as an attack technique for fooling people in phishing and fooling AI systems such as biometric checks.
- Deepfake detection
A defence that checks media for signs it was generated or manipulated by AI, which MITRE ATLAS lists as a mitigation for generating deepfakes, phishing and evading AI models, and whose automated form NIST calls a cat-and-mouse game.
- Defensive distillation
A 2016 defence against adversarial examples that retrains a classifier on its own softened outputs, later shown to hide the gradients attacks use rather than make the model robust.
- Delayed-trigger prompt injection
A prompt injection that plants a dormant rule which only activates later, when a trigger phrase, event or condition appears.
- Denial of AI service
Degrading or shutting down an AI system by exhausting the compute it needs to answer.
- Detecting data poisoning
Whether poisoning can be found depends on what the attacker wanted, not on how advanced the attack was.
- Differential privacy
A mathematical privacy guarantee that an analysis, or a trained model, turns out about the same whether or not any one person's data was used.
- Fast gradient sign method
The one-step attack that builds an adversarial example by nudging every input value the same small amount in whichever direction raises the model's error.
- Gradient masking
When a defence against adversarial examples leaves a model without useful gradients, so gradient-based attacks fail even though the adversarial inputs are still there to be found.
- Hidden context exposure
The unauthorized extraction, inference or reconstruction of the instructions and context an application hides from the user but places in a language model's prompt — OWASP's 2026 category, replacing the narrower System Prompt Leakage.
- Improper output handling
The vulnerability of passing a language model's output to a browser, database, shell or other system without validating, sanitizing or encoding it first.
- Indirect prompt injection
Instructions hidden in third-party content a language model later reads, which the user never sees.
- Jailbreaking
An attack that talks a language model into overriding its own safety training, usually run by the person operating the model rather than a third party.
- Label flipping
Poisoning a model through the labels rather than the samples, which NIST makes cheap for changing one prediction and expensive for degrading the model overall.
- LLM data leakage
An attack in which a crafted prompt gets a language model to reveal sensitive information from its training data, a connected data source, or another user's session.
- LLM prompt self-replication
A prompt injection that makes a language model copy the malicious prompt into its own output, so the prompt spreads to other systems and persists.
- LLM Scope Violation
A named mechanism in which untrusted content directs a model to pull privileged data into its context and act on it without consent.
- LLMjacking
Gaining unauthorized access to a victim's paid language model services, most often with stolen credentials, and using or reselling that access while the victim pays the bill.
- Many-shot jailbreaking
A jailbreak that fills one long prompt with hundreds of faked exchanges of an assistant answering harmful questions, so the model continues the pattern on the real one.
- Membership inference
An attack that asks whether a specific record was in a model's training data, using only the answers the model gives.
- MITRE ATLAS
MITRE's public knowledge base of adversary tactics, techniques and case studies against AI systems.
- Model alignment
Training or fine-tuning a generative AI model so its own behaviour follows chosen values and policies, which MITRE ATLAS lists as a mitigation for prompt injection, jailbreaks, and other generative AI attacks.
- Model evasion
The MITRE ATLAS technique of getting a deployed AI model to misjudge its input at inference time, by crafting adversarial data or presenting a deepfake as genuine.
- Model extraction
Rebuilding a usable copy of a model by querying its API and training on the answers.
- Model inversion
An attack that reconstructs a representative of a model's training data, such as a recognizable face, from the confidence scores its inference API returns.
- Model poisoning
Corrupting a model by tampering with the data or weights it learns from.
- Model provenance
The verifiable history of where a machine learning model came from, meaning the earlier models its weights were derived from and how they were changed on the way.
- OWASP LLM Top 10
OWASP's ranked list of the ten most critical security risks in applications built on large language models. The 2026 edition is the first one checked against a corpus of real incidents.
- Phishing
Social engineering delivered electronically to trick a person into giving up access, credentials or money, which MITRE ATLAS catalogues as an attack technique and says generative AI is helping attackers scale its targeted form.
- Physical adversarial example
An adversarial example engineered to keep working after being printed, worn or projected onto a real object and then seen through a camera, before the model reads it.
- Pickle
Python's built-in format for saving objects and PyTorch's default save format, in which a file is a list of instructions that can import and call functions, so loading an untrusted one can run code.
- Projected gradient descent
The iterative attack that builds an adversarial example in many small steps, each aimed at raising the model's loss and then pulling the input back inside the allowed change.
- Prompt injection
An attack that hides instructions in text a language model reads.
- Prompt injection mitigations
The controls that limit what a fooled language model can do, since no control reliably stops it being fooled.
- Promptware
Malware whose payload is a prompt, not code — text, image or audio that jailbreaks a language model into misusing the application it serves.
- Promptware kill chain
A seven-stage model of an attack on an LLM application, in which prompt injection is the first stage rather than the whole attack.
- Publish poisoned AI artifacts
An attack in which an adversary creates or modifies an AI artifact and publishes it through a public or shared channel for victims to acquire.
- PuzzleMask
A prompt-crafting technique that hides a forbidden instruction inside ordinary prose, so a fast AI safety filter passes it while a stronger model behind the filter decodes and acts on it.
- RAG poisoning
Planting malicious or false content in the data a retrieval-augmented generation system indexes, so a normal query retrieves the attacker's content instead of a true answer.
- Résumé prompt injection
Hidden text a job applicant plants in a résumé so that an AI screener reads it and a human reviewer does not, to push the application up the list.
- Reward hacking
A failure mode where an AI system finds a way to maximize its reward or objective function that technically satisfies the specification while violating the designer's actual intent.
- Role confusion
A failure in which a language model infers who is speaking from how text sounds rather than from the role tag it was given.
- Salami Attack
An automated multi-turn jailbreak that chains many individually harmless-looking prompts, each below a model's refusal threshold, until their combined intent triggers the harmful output a direct request would have been refused for.
- Sanitize training data
Removing an attacker's samples from a training set, a control that works in inverse proportion to how carefully the attacker worked.
- Shadow AI
The use of AI tools, models or AI features inside an organization without the approval or oversight of the people responsible for its IT and security.
- Side-channel model extraction
Recovering a model's architecture or weights by watching the hardware that runs it, rather than relying on its prediction API's answers alone.
- Slopsquatting
A supply chain attack that exploits the predictable tendency of LLMs to hallucinate package names that do not exist, by registering those names on public registries with malicious payloads before anyone else does.
- Subpopulation poisoning
A data poisoning attack that raises a model's error rate on one group of ordinary inputs while barely changing it for everyone else.
- System prompt extraction
An attack that recovers a language model's hidden system prompt, typically by prompting the model itself to reveal it.
- Transferability
The property that an adversarial example made to fool one model often fools a different model trained for the same task.
- Unsafe AI artifacts
Model files and their bundled components, built or altered so that loading or running them harms the victim, by running code on the host or by changing what the model does.
- Verify AI artifacts
Checking the cryptographic checksum or signature of a model file or other AI artifact before using it, so a file an attacker modified is caught instead of loaded.
- Vibe hacking
Using an AI coding agent as an active operator in a cyberattack, so that the agent runs the intrusion rather than only writing code for it.
- Web-scale data poisoning
Poisoning a training set that ships as a list of URLs, by taking control of what those URLs return after the dataset was reviewed.