What matters in AI.

Subscribe

Learn / Category

Learn

AI agents

56 topics in this category.

Guides

2 guides

  • AI agent attacks

    Every named attack on an AI agent, grouped by the step of the agent's loop it targets.

  • How to scope write permissions for LLM agents

    An LLM agent that changes staging or production config should propose changes rather than apply them, hold only short-lived task-scoped credentials, and be checked by a policy the model cannot change.

Definitions, A to Z

54 definitions

  • Agent Control Standard

    An OWASP open specification for runtime governance of AI agents, where a separate Guardian Agent allows, denies, or modifies an agent's actions through middleware hooks.

  • Agent data injection

    An indirect prompt injection that corrupts the trusted metadata an AI agent relies on, so the agent acts on attacker-controlled data while still doing the user's task.

  • Agent harness

    The software layer that wraps a language model, manages its context and tool use, and runs the loop that turns a model into an agent. NIST calls it scaffolding software; Anthropic places it as the third layer in the four-layer framework.

  • Agent hijacking

    An attack that redirects an AI agent by hiding instructions in the data it reads.

  • Agent memory poisoning

    An attack that writes attacker-controlled content into an AI agent's long-term memory, so the agent acts on it in later sessions.

  • Agent skill poisoning

    An attack that hides malicious instructions inside a SKILL.md agent skill file, so the payload runs with the loading agent's own filesystem, credential and network access.

  • Agent-to-agent lateral movement

    An attack in which one compromised AI agent uses its own trusted connections to reach, instruct or impersonate other agents, so the attacker crosses systems without a network path or a stolen credential.

  • Agent2Agent protocol (A2A)

    An open protocol, launched by Google in April 2025 and now a Linux Foundation project, that lets AI agents built by different vendors find each other, hand off tasks and exchange results.

  • AgentDojo

    An open-source test bed that measures whether AI agents can do their job while resisting instructions hidden in the data they read.

  • Agentic identity and privilege abuse

    An AI agent risk in which the credentials, delegated tokens or inherited trust an agent holds are exploited to reach systems and actions beyond the task it was given.

  • Agents Rule of Two

    Meta's rule that an AI agent should hold no more than two of three risky properties in a session.

  • AI agent

    A generative AI model wrapped in software that gives it tools, so its decisions become actions against real systems rather than text on a screen.

  • AI agent context poisoning

    An attack that plants content in what an AI agent treats as trusted context — its saved memory or its running conversation thread — so the change outlives the turn that introduced it.

  • AI agent security

    Protecting the system that lets an AI model act — its tools, harness, identity, memory and environment — rather than the model itself.

  • AI agent tool poisoning

    An attack that corrupts a tool an AI agent already trusts — its description, its code, or the responses it returns — so a routine call carries out the attacker's instructions instead.

  • AI workflow

    A system in which LLMs and tools are orchestrated through predefined code paths — the deterministic end of the agentic spectrum, distinct from agents whose model directs the process.

  • BragJack

    An attack in which a malicious browser extension gets its code onto a page that a browser AI agent trusts, then sends the agent commands or whole prompts of its own.

  • CaMeL

    A prompt injection defense for AI agents that fixes the plan before reading untrusted data and checks every tool call against security policies.

  • Cascading failures

    OWASP's category for a single agentic fault that propagates and amplifies across agents, sessions, or workflows into system-wide harm.

  • Computer-use agent

    An AI agent that acts on a whole desktop through screenshots and coordinates, not just a browser, so its blast radius extends to the file system, the terminal and any installed application.

  • Cross-site prompt injection

    Indirect prompt injection where the page holding the attack and the action it triggers on a web agent are two different places.

  • Dual LLM pattern

    A design for AI assistants that keeps untrusted content away from the model that can use tools, by splitting the work between a privileged and a quarantined model.

  • Excessive agency

    A vulnerability that lets an LLM agent perform actions beyond what its task requires, turning minor malfunctions into serious harm.

  • Four-layer framework

    Anthropic's division of an AI agent into model, tools, harness and execution environment, to argue that the lower layers decide what a failure costs.

  • GhostJacking

    An indirect prompt injection attack that hides instructions inside monitoring logs and alerts so an AI agent reviewing them carries the instructions out.

  • Human-agent trust exploitation

    OWASP's category for attacks that exploit a human's trust in an agent's fluency or authority, rather than any flaw in the agent's own behavior.

  • Human-in-the-loop

    A control that makes an AI agent wait for a person to approve an action before taking it, so an attacker cannot make the agent's tools run automatically.

  • Insecure inter-agent communication

    An AI agent risk in which messages between agents lack authentication, integrity or validation checks, so an attacker can intercept, forge, alter or replay them and steer what the receiving agents do.

  • Instruction privilege escalation

    An attack in which an agent harness rebuilds its context and promotes attacker-controlled text into a higher-privileged instruction role.

  • Least privilege

    The security principle that every program, user or agent should hold only the access its job needs, so a mistake can do only limited damage.

  • Least-Agency

    OWASP's principle that an AI agent should get no more autonomy than its task needs, and none where the task needs no agent.

  • Lethal trifecta

    The three agent capabilities that, held together, let an attacker steal your data.

  • LoopTrap

    A red-teaming attack that corrupts an iterative LLM agent's own judgment of task completion, so it keeps looping well past when the task is actually done.

  • Malicious LLM proxy router

    A third-party API router placed between an AI agent and its model provider that abuses its plaintext access to tamper with tool calls or steal credentials.

  • MCP server

    A program that offers tools, resources and prompts to AI applications through the Model Context Protocol, and one of the places an AI agent's trust can be broken.

  • MCP tool poisoning

    A Model Context Protocol server hides attacker instructions inside a tool's description or its responses, instructions the AI agent reads and follows but the person who approved the tool never sees.

  • MCP tool shadowing

    A rogue MCP server registers a tool with a name or description close enough to a trusted one that the agent picks the impostor at selection time.

  • Memory provenance

    Labelling every entry an AI agent writes to memory with where it came from, so a poisoned record can be traced to its source, contained and rolled back.

  • Model Context Protocol (MCP)

    An open standard, released by Anthropic in 2024 and now hosted by the Linux Foundation's Agentic AI Foundation, for connecting AI applications to outside tools and data through MCP servers.

  • ObliInjection

    A prompt-injection attack for LLM agents that succeeds without knowing where the attacker's segment lands among other sources in the input.

  • Oracle poisoning

    Corrupting the structured knowledge graph an AI agent queries at runtime, so it reasons correctly over false facts.

  • OWASP Agentic Skills Top 10

    OWASP's early-stage list of the ten most critical security risks in agent skills, the packages of instructions and scripts that AI agents load and run.

  • OWASP Top 10 for Agentic Applications

    OWASP's list of the ten most critical security risks in AI agents, the systems that plan, call tools, keep memory and act across several steps. Its entries are numbered ASI01 to ASI10.

  • Prompt Forcing

    An attack in which a lower-trust component, such as a browser extension or a website, makes an AI agent accept its prompts as if the user had sent them.

  • Rogue agents

    OWASP's category for AI agents that act harmfully, deceptively, or parasitically after their behavior has already drifted from its authorized scope.

  • STAC (sequential tool attack chaining)

    A multi-turn attack that chains individually harmless tool calls into a harmful outcome, so no single call in the sequence ever looks malicious.

  • Structured template injection

    An agent-hijacking attack that hides forged chat-template turn markers in content an AI agent retrieves, so attacker text reads as a user request.

  • Tool misuse

    An AI agent risk in which an agent uses a legitimate tool it is allowed to use in an unsafe or unintended way, such as deleting data, running up API costs or sending information out.

  • ToolHijacker

    A prompt injection attack that publishes one optimised tool document so an LLM agent's retriever shortlists it and its model picks it for an attacker-chosen task.

  • Unexpected code execution

    OWASP's category for agentic code-generation and execution features turned into remote code execution, host compromise, or sandbox escape.

  • WARP (web agent retrieval poisoning)

    An attack that appends a short persuasive passage to a user-generated page deep-research agents already retrieve, so their reports cite and promote an attacker-chosen entity across a whole topic.

  • Web agent

    An AI agent scoped to a browser — it reads a web page's structure or a screenshot and acts on it by clicking, typing and navigating, distinct from the broader agent that also drives a desktop.

  • WebInject

    A prompt injection attack on screenshot-reading web agents that hides a small, optimised pixel perturbation in a webpage so the agent performs an action the attacker chose.

  • Workflow identity hijacking

    An attack in which a benign request to an unauthenticated intake triggers an AI workflow that acts with its own privileges rather than the requester's, so the model needs no manipulation.