Definition · AI agents
Human-in-the-loop
Human-in-the-loop, in AI agent security, is a control that requires a person to approve an agent's action before the agent carries it out. MITRE ATLAS lists human-in-the-loop as mitigation AML.M0029 against attackers invoking an agent's tools, including to leak or destroy data. The human approver, not the model, makes the final decision.
Last reviewed
Key points
- In agent security, human-in-the-loop means a person approves an agent's action before the agent takes it. MITRE ATLAS lists it as mitigation AML.M0029.
- ATLAS maps it to three techniques, all attacker tool calls through the agent, including leaking and destroying data. It gates the action, not the injection behind it.
- Approval should scale with consequence. ATLAS says minor tasks may need little oversight and high-consequence systems several approvers; OWASP lets reversible actions auto-approve.
- The human becomes the target. OWASP describes hijacked agents writing convincing reasons to get harmful actions approved.
- IBM's explainer uses the term for any human role in an AI workflow, including labelling training data. That role gives none of this protection.
How it works
The agent proposes an action and stops. A person approves or refuses it. A refused action does not run.
MITRE ATLAS sets the rule in mitigation AML.M0029: systems “should require the user or another human stakeholder to approve AI agent actions before the agent takes them”. Tools such as audit agents may help, but “final adjudication should be conducted by a human decision-maker”.
The OWASP GenAI Security Project’s LLM03:2026 entry on excessive agency says the check can sit in the tool or downstream.
ATLAS says the policy “should follow the degree of consequence of the task at hand”. LLM03 gives a refund example: store credit can be undone, so it auto-approves, while an external payout goes to a human.
Why it matters
Agent hijacking lets an attacker use an AI agent’s tools. ATLAS maps AML.M0029 to three techniques: invoking the agent’s tools, and using them to leak or destroy data. For all three, confirmation “can prevent the automatic execution of tools by an adversary”.
In LLM03’s example, a crafted email tells an assistant to forward sensitive mail to the attacker. If the user must “manually review and hit ‘send’ on every mail”, the forward never leaves on its own.
What it does not stop
Human-in-the-loop gates the action, not the attack. The hijacking still happens: the agent proposes what the attacker wants, and the control works only if the person says no.
That makes the person the target. OWASP’s Top 10 for Agentic Applications, in its human-agent trust exploitation entry, lists “Fake Explainability”, where the agent “fabricates convincing rationales” so people approve unsafe actions. Another of its scenarios ends with an analyst approving deletion of a live production database.
The gate can also miss a path. Under the same tool-invocation technique that AML.M0029 mitigates, ATLAS notes that an agent may open an app through deep links or redirect chains even when that app “was not selected as a direct agent tool”, “potentially avoiding normal user-confirmation workflows”.
Trade-offs
Every approval costs a person’s time. ATLAS says the security benefit “may be at odds with operational overhead costs of additional approvals”. Meta’s Agents Rule of Two, in a hypothetical internal coding agent, calls human-in-the-loop “a valuable defense-in-depth” but says developers “aim to unlock operation at scale by minimizing human interventions”. It also names “a background process where human-in-the-loop is disruptive or ineffective” as a hard case.
Where definitions disagree
How wide the term is. IBM’s explainer uses human-in-the-loop for any human involvement “at some point in the AI workflow”, to ensure “accuracy, safety, accountability or ethical decision-making”. That includes approving or overriding outputs, but also labelling training data and training from human feedback. ATLAS uses the term for one control: approving an agent’s action before it runs. A team that only labels its data is doing human-in-the-loop in IBM’s sense and has none of the protection ATLAS means.
Whether the human can be replaced. ATLAS lets audit agents assist, but keeps the final decision with a human. Meta’s rule is that an agent should combine no more than two of three properties in one session: processing untrustworthy inputs, accessing sensitive systems or private data, and changing state or communicating externally. An agent that needs all three requires supervision, and Meta accepts “human-in-the-loop approval or another reliable means of validation”. Meta calls software guardrails and human approvals the preferred methods today, and says it is researching oversight agents to satisfy the same checks.
Questions and answers
What does human-in-the-loop mean for AI agents?
For AI agents, human-in-the-loop means the agent must get a person's approval before it carries out an action, such as sending an email or deleting a file. MITRE ATLAS lists the control as AML.M0029. Audit tools may help the approver, but ATLAS says "final adjudication should be conducted by a human decision-maker".
Does human-in-the-loop stop prompt injection?
No. Human-in-the-loop does not stop an agent from reading and following injected instructions. It puts a person between the agent and the action. ATLAS maps AML.M0029 to tool invocation techniques, where confirmation "can prevent the automatic execution of tools by an adversary". A person who approves the hijacked action still lets the attack through.
Should every agent action need human approval?
No. ATLAS says approval should follow the degree of consequence: minimal oversight for minor, repetitive tasks on basic tools, and possibly approvers from several organizations for high-consequence systems. OWASP's LLM03 describes auto-approving reversible actions, such as a refund as store credit, and routing irreversible ones, such as an external payout, to a person.
Sources
- MITRE ATLAS, AML.M0029 Human In-the-Loop for AI Agent Actions (collection 2026.09)MITRE
- LLM03:2026 Excessive AgencyOWASP GenAI Security Project, 4 Aug 2026
- OWASP Top 10 for Agentic Applications 2026, ASI09: Human-Agent Trust ExploitationOWASP GenAI Security Project, 9 Dec 2025
- Agents Rule of Two: A Practical Approach to AI Agent SecurityMeta, 31 Oct 2025
- What is human-in-the-loop?IBM, 8 Jul 2025