What matters in AI.

Subscribe

Learn / AI agents

Definition · AI agents

AI agent security

AI agent security is the practice of protecting the system an AI agent runs inside, not only the model: the tools it can call, the harness that orchestrates multi-step actions, its identity and permissions, its memory, and the environment those actions land in. AI security protects the model; AI agent security protects what the model is allowed to do.

Last reviewed

Key points

  • AI agent security protects the system around a model that can act — its tools, the harness orchestrating multi-step actions, its identity and permissions, its memory and its environment — not the model's training data or weights, which AI security covers.
  • OWASP's Top 10 for Agentic Applications, December 2025, ranks ten agent-specific risks, ASI01 through ASI10, on the grounds that agents, unlike task-specific automations, plan, decide and act across multiple steps and systems, often on a user's behalf.
  • NIST launched an AI Agent Standards Initiative in February 2026 whose research pillar covers agent authentication and identity infrastructure, alongside pillars for industry-led standards and open-source agent protocols.
  • Anthropic makes the boundary explicit: agent security is 'a property of the whole system, not just the model,' and argues the harness and execution environment decide what a failure costs, not the model.

Each attack against an agent has its own topic here: Agent hijacking, AI agent tool poisoning, AI agent context poisoning and AI supply chain compromise.

Where the boundary sits

An AI agent is a model plus scaffolding that lets it call tools and act across multiple steps, “often on behalf of users and teams,” OWASP says — unlike a “task-specific automation.” That scaffolding is what AI security does not cover: it protects the model’s training data, weights and input handling, not what the model is subsequently permitted to touch.

NIST’s own AI agent work draws the same line by exclusion. Its request for information on agent security scopes itself to systems that can “affect external state,” setting aside chatbots and retrieval systems “not orchestrated to act autonomously.” A sound model wrapped in a system with wide tool access, standing credentials and no oversight is where AI agent security starts.

Why it matters

The consequence, not the mechanism, is what changes. Anthropic’s four-layer framework makes this its central claim: a model failure “is identical” whether or not the agent runs sandboxed, but “the consequences are not,” because the harness and execution environment decide the blast radius, not the model.

That is also why NIST’s AI Agent Standards Initiative, launched in February 2026, puts “agent authentication and identity infrastructure” alongside its standards and protocol work rather than folding agents into existing model guidance. And it is why OWASP, introducing its ten agent risks, extends least privilege into “Least-Agency”: grant an agent no more autonomy than the task needs, because agentic behaviour where it is not needed “expands the attack surface without adding value.”

Where definitions disagree

NIST and OWASP draw the boundary at different widths. NIST’s request for information scopes AI agent security to systems whose actions change state outside themselves — a narrower line that excludes an agent that only drafts or retrieves. OWASP’s Top 10 for Agentic Applications is wider: its first risk, Agent Goal Hijack, counts an agent manipulated into producing “fraudulent information that impacts business decisions,” harm that reaches the world through a person reading the output rather than through an action the agent takes.

Anthropic’s four-layer framework is not a competing scope so much as an unsettled one. It is a proposal made in a March 2026 comment to NIST, not adopted guidance, and its own document does not hold to four layers consistently — a “product layer” appears alongside the model, tools, harness and environment it names as the four.

Questions and answers

What is AI agent security?

AI agent security is the practice of protecting the system an AI agent runs inside — the tools it can call, the harness that orchestrates multi-step actions, its identity and permissions, its memory, and the environment those actions land in — rather than the model alone.

How is AI agent security different from AI security?

AI security protects the model: its training data, its weights, and how it reads its input. AI agent security protects what a model is allowed to do once it can call tools and act across multiple steps — a model can be sound and the system built around it still unsafe, because the failure that matters is what the agent's tools and permissions let a bad decision reach.

What frameworks define AI agent security risks?

OWASP's Top 10 for Agentic Applications 2026 names ten risk categories, ASI01 through ASI10, covering goal hijacking, tool misuse, identity and privilege abuse, supply chain, code execution, memory poisoning, inter-agent communication, cascading failures, human-agent trust and rogue agents. NIST's AI Agent Standards Initiative, launched February 2026, conducts research into agent authentication and identity rather than publishing its own risk taxonomy. Anthropic's four-layer framework is a third account, proposed rather than standardised.

Sources

  1. OWASP Top 10 for Agentic Applications 2026OWASP Gen AI Security Project, 9 Dec 2025
  2. Announcing the "AI Agent Standards Initiative" for Interoperable and Secure InnovationNIST, 17 Feb 2026
  3. AI Agent Standards InitiativeNIST, 17 Feb 2026
  4. Request for Information Regarding Security Considerations for Artificial Intelligence AgentsCenter for AI Standards and Innovation, NIST, US Department of Commerce, 8 Jan 2026
  5. Re: Request for Information: Security Considerations for Artificial Intelligence Agents, Docket No. NIST-2025-0035Anthropic, 9 Mar 2026

Guides that use this term