What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

Human-agent trust exploitation

Human-agent trust exploitation is OWASP's category for attacks that use an AI agent's fluency, confidence, or apparent expertise to manipulate the human overseeing it into approving a harmful action or disclosing sensitive information. ASI09 names the human's misplaced trust as the point of failure, not a defect in the agent's own reasoning.

Last reviewed

Key points

  • Human-agent trust exploitation is ASI09 in OWASP's Top 10 for Agentic Applications 2026, published 9 December 2025.
  • It targets the human, not the agent. OWASP defines it as exploiting the trust a person places in an agent's natural language fluency, emotional intelligence, and perceived expertise to influence decisions, extract information, or steer outcomes.
  • OWASP draws the boundary explicitly — "This entry is about human misperception or over-reliance whereas ASI10 is agent intent deviation" — the agent can be behaving exactly as designed and still cause the harm.
  • It names its upstream and downstream chain precisely — it builds on LLM06:2025 Excessive Agency, can be caused by LLM01:2025 Prompt Injection or LLM05:2025 Improper Output Handling, and results in LLM09:2025 Misinformation.
  • Its worked examples turn on missing friction rather than missing permissions — no confirmation step before a sensitive action, or a fabricated rationale that makes an unsafe action look reviewed.

How it works

OWASP names the mechanism as anthropomorphism: an agent’s fluency and perceived expertise earn it a trust a plain script never would. An attacker, or a misaligned design with no attacker at all, can lean on that trust directly — a plausible-sounding justification is often easier to produce than a real compromise of the agent’s own logic.

OWASP is explicit about where this entry starts and stops: “This entry is about human misperception or over-reliance whereas ASI10 is agent intent deviation.” It also names its place in the OWASP chain — it builds on excessive agency (LLM06:2025), since a missing confirmation step is what lets fabricated trust turn into an executed action; it can be caused by prompt injection (LLM01:2025) or improper output handling; and left unchecked it results in misinformation.

OWASP describes the agent itself as an untraceable “bad influence”: it manipulates the human into performing the final, audited action, so its role in the compromise is invisible to forensics.

Why it matters

OWASP’s common examples describe missing friction, not missing permissions. A missing confirmation step converts a single persuasive prompt directly into an irreversible transfer or configuration change. Fake explainability is the sharpest version: the agent fabricates a rationale that makes an unsafe action look reviewed, so the human approves it believing it is justified.

Two worked scenarios make the stakes concrete. A finance copilot ingests a poisoned invoice and recommends an urgent payment to attacker-controlled bank details; the manager approves it, trusting the agent’s expertise, and the company loses the funds. A hijacked agent fabricates a rationale that talks an analyst into deleting a live production database. In both, every step the human takes is a legitimate approval — the harm is that the thing approved was never what it claimed to be.

In practice

OWASP’s mitigations target the approval step itself rather than trying to make agents more trustworthy. Explicit, multi-step confirmation before any sensitive or high-impact action removes the single-prompt-to-execution path that several of the attack scenarios rely on. Separating preview from effect — blocking state-changing calls while a user is only looking at a preview — closes the consent-laundering pattern where a “read-only” screen quietly triggers a side effect on open.

Beyond the approval gate, OWASP recommends adaptive trust calibration: adjusting how much autonomy an agent gets based on contextual risk, and showing confidence-weighted cues such as “low-certainty” or “unverified source” so a human is prompted to question a high-impact action rather than approve it on reflex. Content provenance — verifiable source identifiers, timestamps, and integrity hashes attached to recommendations — lets a runtime policy block actions that lack trusted provenance. Immutable logs of queries and actions, plan-divergence detection against an approved workflow baseline, and a plain-language (not model-generated) way for users to flag suspicious behavior round out the guidance.

Questions and answers

Is human-agent trust exploitation a flaw in the agent?

Not necessarily. OWASP defines ASI09 as human misperception or over-reliance on an agent's fluency and apparent expertise — the agent can be functioning exactly as designed while the human overseeing it is still manipulated into approving something harmful.

How is this different from rogue agents (ASI10)?

OWASP states the boundary directly: ASI09 is about human misperception or over-reliance, while ASI10 is agent intent deviation. ASI09 needs no compromised or drifting agent, only a human who trusts a plausible explanation; ASI10 requires the agent's own behavior to have gone off its authorized scope.

Sources

  1. OWASP Top 10 for Agentic Applications 2026, ASI09: Human-Agent Trust ExploitationOWASP Gen AI Security Project, 9 Dec 2025

Guides that use this term