Definition · AI basics
AI safety
AI safety is the work of keeping AI systems from endangering human life, health, property or the environment, including through errors and unexpected behaviour when nobody is attacking them. NIST lists safety as a separate characteristic from security and resilience, though the two overlap, most clearly when an attacker defeats a system's safety controls.
Last reviewed
Key points
- AI safety asks whether an AI system could endanger people, property or the environment. AI security asks whether someone could attack it.
- NIST lists "safe" and "secure and resilient" as separate characteristics of trustworthy AI. A safe system should not, "under defined conditions," endanger "human life, health, property, or the environment."
- This site's working test for the boundary is whether there is an attacker. It is a rule of thumb, not NIST's. NIST defines safety by the harm, and counts recovering from unexpected adverse events as resilience, part of security.
- The clearest overlap is an attacker defeating a safety control, which NIST files under adversarial machine learning. NIST says that field is "neither a complete solution to, nor a subset of" AI safety.
- Deliberate misuse carries both labels. The International AI Safety Report lists it among its risks, and the UK renamed its AI Safety Institute the AI Security Institute with a focus on it.
The adversary test
AI safety asks whether a system could endanger people, property or the environment. AI security asks whether someone could attack it. The working test used on this site is whether there is an attacker. A dangerous failure with no attacker is a safety problem. An attack that endangers people is both.
The test is a rule of thumb, not NIST’s wording. NIST’s AI RMF 1.0 lists “safe” and “secure and resilient” as separate characteristics of trustworthy AI, and defines safety by the harm, not its cause. AI systems should “not under defined conditions, lead to a state in which human life, health, property, or the environment is endangered.”
NIST’s security covers attacks such as data poisoning, but it also includes resilience, “the ability to return to normal function after an unexpected adverse event,” attacker or not.
What falls on each side
A model that is simply wrong, as in AI hallucination, has no attacker. NIST puts failures “outside the context of adversarial use, such as inaccuracy” outside adversarial machine learning. Such a failure is a safety problem when it endangers someone.
The clearest overlap is an attacker defeating a safety control. NIST calls this goal “misuse enablement”: circumventing restrictions the owner places on a system, such as restrictions “designed to prevent the system from producing outputs that could cause harm to others.” A jailbreak aimed at those restrictions is a security attack on a safety control.
Defending against attacks helps safety without covering it. NIST says robustness to these attacks “may play an important role” in AI safety.
Where definitions disagree
Misuse is where the adversary test runs out. A person who uses a model as designed to plan a crime is an adversary, but is not attacking the model.
The International AI Safety Report 2026 puts misuse first among its three risk categories: “Risks from misuse, where actors deliberately use AI systems to cause harm.” The report says the categories are “not exhaustive or mutually exclusive.”
The UK government calls the same ground security. On 14 February 2025 it renamed its AI Safety Institute the AI Security Institute. The new name reflects a focus on risks such as “how the technology can be used to develop chemical and biological weapons, how it can be used to carry out cyber-attacks, and enable crimes such as fraud and child sexual abuse.” Those examples are people using AI to cause harm, not attacks on the model.
The difference is mostly the label. The UK AI Security Institute is part of the safety report’s secretariat, and the government said the institute’s work “won’t change.”
Questions and answers
Is AI safety the same as AI security?
No, though they overlap. AI safety asks whether a system could endanger people, property or the environment, whatever the cause. AI security asks whether someone could attack it. NIST's AI RMF lists "safe" and "secure and resilient" as separate characteristics of trustworthy AI, and NIST's adversarial machine learning report says that field is "neither a complete solution to, nor a subset of" AI safety.
Is a jailbreak a safety problem or a security problem?
A jailbreak is a security problem, because it is a deliberate attack. It is also a safety problem when the control it defeats exists to prevent harm. NIST calls the attacker's goal "misuse enablement": circumventing restrictions the owner places on a system, such as restrictions "designed to prevent the system from producing outputs that could cause harm to others," and files the techniques under adversarial machine learning.
Is a hallucination an AI safety problem?
A hallucination is a safety problem when the wrong answer endangers life, health, property or the environment. With no attacker involved it falls outside adversarial machine learning: NIST places failures "outside the context of adversarial use, such as inaccuracy" outside that field.
Sources
- Artificial Intelligence Risk Management Framework (NIST AI 100-1, AI RMF 1.0)NIST, 26 Jan 2023
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (NIST AI 100-2e2025)NIST, 24 Mar 2025
- International AI Safety Report 2026UK Department for Science, Innovation and Technology (DSIT 2026/001), 3 Feb 2026
- Tackling AI security risks to unleash growth and deliver Plan for ChangeUK Department for Science, Innovation and Technology, 14 Feb 2025