What matters in AI.

Subscribe

Learn / Category

Learn

AI security

87 topics in this category.

Guides

9 guides

  • Guardrail bypass techniques

    The named ways attackers get a language model past its safety controls, grouped by the control each one mainly gets around, plus the reconnaissance that comes first.

  • How to detect shadow AI

    Shadow AI can be found in proxy and firewall logs, identity-provider app consents, code repositories, cloud audit logs and scans for self-hosted models. Each source misses something, so check several and give staff a safe way to report what they use.

  • How to red team an LLM application

    Red team an LLM application by starting from the harm that would matter, attacking the whole deployed application where untrusted text gets in, probing by hand before automating, and repeating every attack.

  • Input filtering vs output validation

    Input filtering screens what goes into an LLM; output validation checks what comes out before a user or another system acts on it. Each misses what the other catches, and OWASP and MITRE ATLAS list both.

  • Security risks of running LLMs locally

    Running a language model locally keeps prompts away from a provider, but it puts a downloaded model file, an inference server and the model's behaviour on your own machine, and it leaves prompt injection and hallucinated output in place.

  • Shadow AI vs shadow IT

    Shadow AI is the part of shadow IT that involves AI tools. The shadow IT playbook still applies, but AI adds three things: prompts may train the provider's models depending on the account, AI can arrive through an update to software already in use, and the output itself can be wrong.

  • State of AI security, September 2026 edition

    Seven pillars of AI security, ranked by where problems are actually being found rather than by where the noise is.

  • Targeted vs untargeted data poisoning

    Untargeted poisoning aims to make a model worse on inputs in general. Targeted poisoning changes its answer on a few inputs the attacker picked and leaves the rest alone, which is why accuracy checks usually catch the first and can miss the second.

  • White-box vs black-box attacks

    A white-box attacker knows the model inside out; a black-box attacker knows little or nothing about it and usually works by sending inputs and reading outputs. The labels describe what the attacker knows, not what the attacker does, and NIST describes a gray-box range in between.

Definitions, A to Z

78 definitions

  • Adaptive attack

    An attack designed against one specific defence by an attacker who knows how it works, and the standard test of defences against adversarial examples.

  • Adversarial example

    An input crafted so a machine learning model gets it wrong at deployment time.

  • Adversarial machine learning

    The field that studies attacks exploiting how machine learning works, and how to build models that withstand them.

  • Adversarial suffix

    A string of tokens, found by automated search, that is appended to a harmful request so an aligned language model answers instead of refusing; the best-known method for finding one is GCG.

  • Adversarial training

    A defence that trains a model on the adversarial examples an attacker would use against it, so it learns to get them right.

  • AI bill of materials

    A machine readable inventory of the datasets, models and components a system depends on.

  • AI dataset provenance

    A maintained history of where a training dataset came from and every change made to it since.

  • AI guardrails

    Safety controls placed between users, tools, and a generative AI model that check prompts, retrieved content, outputs, and agent actions before they take effect.

  • AI output watermarking

    Hiding a machine-detectable signal in what an AI model generates, so the output can later be traced to that model, and why the signal can be removed.

  • AI red teaming

    Authorised adversary-emulation exercises run against a deployed AI system to find its weaknesses before an attacker does.

  • AI security

    Protecting machine learning systems from attack, including attacks on the model itself.

  • AI security posture management

    A set of vendor security tools that discover, monitor, and remediate AI-specific risk — exposed model weights, over-permissioned AI service accounts, exposed training data — across an organization's AI systems.

  • AI security trade-offs

    The costs an organisation accepts when it hardens an AI system against attack, and the decision NIST says the organisation must then make.

  • AI supply chain compromise

    An attack that reaches an AI system through a part it acquired — a model, a dataset, a framework, a container image or an agent tool.

  • AI supply chain rug pull

    An attack where an AI component is published as legitimate, gains adoption, and is then updated to a malicious version, which may let the adversary bypass the extra scrutiny new components often receive.

  • Architectural backdoor

    A backdoor built into a model's structure rather than its weights or training data, and in some designs it survives retraining.

  • ASCII smuggling

    Hiding text inside invisible Unicode tag characters, so a human sees nothing while a language model reads the hidden message.

  • Attack success rate

    The share of attack attempts in a test that achieved the attacker's goal, and why it describes the test as much as the system.

  • Backdoor attack

    A hidden conditional rule inside a trained model, dormant until an input carries the attacker's trigger.

  • Clean-label poisoning

    Poisoning a training set without touching a single label, which removes the mismatch every cheap filter is looking for.

  • Confused deputy

    A program tricked into using its own authority for someone who should not have it, because it cannot tell which of its principals a request comes from.

  • Cost harvesting

    Running up the bill for a victim's paid AI service by driving it past its normal operating capacity.

  • Crescendo

    A multi-turn jailbreak that opens with a harmless question about a topic and escalates a step at a time, each new prompt building on what the model itself just wrote, until the model produces content it would have refused if asked directly.

  • Data poisoning

    Controlling part of the data a model trains on, so the model learns what the attacker chose.

  • Deepfake

    Synthetic media made with generative AI to pass as authentic, which MITRE ATLAS catalogues as an attack technique for fooling people in phishing and fooling AI systems such as biometric checks.

  • Deepfake detection

    A defence that checks media for signs it was generated or manipulated by AI, which MITRE ATLAS lists as a mitigation for generating deepfakes, phishing and evading AI models, and whose automated form NIST calls a cat-and-mouse game.

  • Defensive distillation

    A 2016 defence against adversarial examples that retrains a classifier on its own softened outputs, later shown to hide the gradients attacks use rather than make the model robust.

  • Delayed-trigger prompt injection

    A prompt injection that plants a dormant rule which only activates later, when a trigger phrase, event or condition appears.

  • Denial of AI service

    Degrading or shutting down an AI system by exhausting the compute it needs to answer.

  • Detecting data poisoning

    Whether poisoning can be found depends on what the attacker wanted, not on how advanced the attack was.

  • Differential privacy

    A mathematical privacy guarantee that an analysis, or a trained model, turns out about the same whether or not any one person's data was used.

  • Fast gradient sign method

    The one-step attack that builds an adversarial example by nudging every input value the same small amount in whichever direction raises the model's error.

  • Gradient masking

    When a defence against adversarial examples leaves a model without useful gradients, so gradient-based attacks fail even though the adversarial inputs are still there to be found.

  • Hidden context exposure

    The unauthorized extraction, inference or reconstruction of the instructions and context an application hides from the user but places in a language model's prompt — OWASP's 2026 category, replacing the narrower System Prompt Leakage.

  • Improper output handling

    The vulnerability of passing a language model's output to a browser, database, shell or other system without validating, sanitizing or encoding it first.

  • Indirect prompt injection

    Instructions hidden in third-party content a language model later reads, which the user never sees.

  • Jailbreaking

    An attack that talks a language model into overriding its own safety training, usually run by the person operating the model rather than a third party.

  • Label flipping

    Poisoning a model through the labels rather than the samples, which NIST makes cheap for changing one prediction and expensive for degrading the model overall.

  • LLM data leakage

    An attack in which a crafted prompt gets a language model to reveal sensitive information from its training data, a connected data source, or another user's session.

  • LLM prompt self-replication

    A prompt injection that makes a language model copy the malicious prompt into its own output, so the prompt spreads to other systems and persists.

  • LLM Scope Violation

    A named mechanism in which untrusted content directs a model to pull privileged data into its context and act on it without consent.

  • LLMjacking

    Gaining unauthorized access to a victim's paid language model services, most often with stolen credentials, and using or reselling that access while the victim pays the bill.

  • Many-shot jailbreaking

    A jailbreak that fills one long prompt with hundreds of faked exchanges of an assistant answering harmful questions, so the model continues the pattern on the real one.

  • Membership inference

    An attack that asks whether a specific record was in a model's training data, using only the answers the model gives.

  • MITRE ATLAS

    MITRE's public knowledge base of adversary tactics, techniques and case studies against AI systems.

  • Model alignment

    Training or fine-tuning a generative AI model so its own behaviour follows chosen values and policies, which MITRE ATLAS lists as a mitigation for prompt injection, jailbreaks, and other generative AI attacks.

  • Model evasion

    The MITRE ATLAS technique of getting a deployed AI model to misjudge its input at inference time, by crafting adversarial data or presenting a deepfake as genuine.

  • Model extraction

    Rebuilding a usable copy of a model by querying its API and training on the answers.

  • Model inversion

    An attack that reconstructs a representative of a model's training data, such as a recognizable face, from the confidence scores its inference API returns.

  • Model poisoning

    Corrupting a model by tampering with the data or weights it learns from.

  • Model provenance

    The verifiable history of where a machine learning model came from, meaning the earlier models its weights were derived from and how they were changed on the way.

  • OWASP LLM Top 10

    OWASP's ranked list of the ten most critical security risks in applications built on large language models. The 2026 edition is the first one checked against a corpus of real incidents.

  • Phishing

    Social engineering delivered electronically to trick a person into giving up access, credentials or money, which MITRE ATLAS catalogues as an attack technique and says generative AI is helping attackers scale its targeted form.

  • Physical adversarial example

    An adversarial example engineered to keep working after being printed, worn or projected onto a real object and then seen through a camera, before the model reads it.

  • Pickle

    Python's built-in format for saving objects and PyTorch's default save format, in which a file is a list of instructions that can import and call functions, so loading an untrusted one can run code.

  • Projected gradient descent

    The iterative attack that builds an adversarial example in many small steps, each aimed at raising the model's loss and then pulling the input back inside the allowed change.

  • Prompt injection

    An attack that hides instructions in text a language model reads.

  • Prompt injection mitigations

    The controls that limit what a fooled language model can do, since no control reliably stops it being fooled.

  • Promptware

    Malware whose payload is a prompt, not code — text, image or audio that jailbreaks a language model into misusing the application it serves.

  • Promptware kill chain

    A seven-stage model of an attack on an LLM application, in which prompt injection is the first stage rather than the whole attack.

  • Publish poisoned AI artifacts

    An attack in which an adversary creates or modifies an AI artifact and publishes it through a public or shared channel for victims to acquire.

  • PuzzleMask

    A prompt-crafting technique that hides a forbidden instruction inside ordinary prose, so a fast AI safety filter passes it while a stronger model behind the filter decodes and acts on it.

  • RAG poisoning

    Planting malicious or false content in the data a retrieval-augmented generation system indexes, so a normal query retrieves the attacker's content instead of a true answer.

  • Résumé prompt injection

    Hidden text a job applicant plants in a résumé so that an AI screener reads it and a human reviewer does not, to push the application up the list.

  • Reward hacking

    A failure mode where an AI system finds a way to maximize its reward or objective function that technically satisfies the specification while violating the designer's actual intent.

  • Role confusion

    A failure in which a language model infers who is speaking from how text sounds rather than from the role tag it was given.

  • Salami Attack

    An automated multi-turn jailbreak that chains many individually harmless-looking prompts, each below a model's refusal threshold, until their combined intent triggers the harmful output a direct request would have been refused for.

  • Sanitize training data

    Removing an attacker's samples from a training set, a control that works in inverse proportion to how carefully the attacker worked.

  • Shadow AI

    The use of AI tools, models or AI features inside an organization without the approval or oversight of the people responsible for its IT and security.

  • Side-channel model extraction

    Recovering a model's architecture or weights by watching the hardware that runs it, rather than relying on its prediction API's answers alone.

  • Slopsquatting

    A supply chain attack that exploits the predictable tendency of LLMs to hallucinate package names that do not exist, by registering those names on public registries with malicious payloads before anyone else does.

  • Subpopulation poisoning

    A data poisoning attack that raises a model's error rate on one group of ordinary inputs while barely changing it for everyone else.

  • System prompt extraction

    An attack that recovers a language model's hidden system prompt, typically by prompting the model itself to reveal it.

  • Transferability

    The property that an adversarial example made to fool one model often fools a different model trained for the same task.

  • Unsafe AI artifacts

    Model files and their bundled components, built or altered so that loading or running them harms the victim, by running code on the host or by changing what the model does.

  • Verify AI artifacts

    Checking the cryptographic checksum or signature of a model file or other AI artifact before using it, so a file an attacker modified is caught instead of loaded.

  • Vibe hacking

    Using an AI coding agent as an active operator in a cyberattack, so that the agent runs the intrusion rather than only writing code for it.

  • Web-scale data poisoning

    Poisoning a training set that ships as a list of URLs, by taking control of what those URLs return after the dataset was reviewed.