Definition · AI agents
AI agent
An AI agent is a system in which a generative AI model decides what to do and surrounding software carries it out. The model emits tool calls; scaffolding software executes them against real systems, with real permissions. That is what separates an AI agent from a chat application — the agent acts, rather than only answering.
Last reviewed
Key points
- An AI agent is a generative AI model plus scaffolding software — at least one model, NIST says, and scaffolding that equips it with tools to take discretionary actions.
- What separates it from a chat application is acting. NIST asks only about agents whose actions change state outside themselves, setting aside chatbots and retrieval systems not orchestrated to act autonomously.
- The model never executes anything. It emits a tool call, and the scaffolding runs it with whatever permissions the deployment granted.
- Membership is contested. Anthropic calls a predefined code path a workflow rather than an agent; OpenAI's governance paper says there is no clear line to draw.
- The older definitions require no language model. OWASP defines an agent without naming one, and credits the OpenAI Gym with helping drive the first wave of agentic AI.
An AI agent is not a smarter chatbot. The model produces a tool call, and software outside the model reads it and acts on something real.
What makes it an agent and not a chat application
NIST’s Center for AI Standards and Innovation characterised agent systems in a January 2026 request for information. “AI agent systems consist of at least one generative AI model and scaffolding software that equips the model with tools to take a range of discretionary actions.” They are “capable of planning and taking autonomous actions that impact real-world systems or environments”.
The notice’s scope is narrower than that characterisation. It asks only about agents “capable of taking actions that affect external state, i.e., persistent changes outside of the AI agent system itself”, and sets aside “AI chatbots or retrieval-augmented generation systems that are not orchestrated to act autonomously”. What that puts outside is not the chat window but the absence of orchestration to act.
Why the split between deciding and acting matters
The model never executes anything. Between decision and action sits the harness — Anthropic’s name for “the wrapper that lets a model run as an agent at all”, which “manages prompts, formats tool calls, carries context across turns”, and runs the loop. Whatever permissions that software holds are the ones a wrong decision spends.
Checks exist — NIST names instruction hierarchy as one — but a hierarchy is a trained tendency, not authentication. Nothing in the loop can verify where an instruction came from, so a model obeying the hierarchy faithfully can still be obeying the attacker. That is why prompt injection against an agent becomes agent hijacking — a sent email rather than a bad paragraph — and why Anthropic’s four-layer framework argues the layers beneath the model decide what a failure costs.
Where definitions disagree
Three things are genuinely unsettled, and each changes what a claim about “agents” covers.
Whether it is a category or a degree. OpenAI’s own governance white paper refuses the line, saying there is “no clear line along which to draw a binary distinction” between agents and current AI systems like GPT-4. It defines the degree of agenticness in a system as “the degree to which a system can adaptably achieve complex goals in complex environments with limited direct supervision”, and prefers the term agentic AI systems because agenticness is “a property rather than a category/classification”. Chan and colleagues do the same for algorithmic systems generally, naming four characteristics that come in degrees: underspecification, directness of impact, goal-directedness and long-term planning. OpenAI’s later practical guide draws a line anyway: “Applications that integrate LLMs but don’t use them to control workflow execution—think simple chatbots, single-turn LLMs, or sentiment classifiers—are not agents.”
Who has to control the loop. Anthropic splits agentic systems in two. Workflows are “systems where LLMs and tools are orchestrated through predefined code paths”; agents are “systems where LLMs dynamically direct their own processes and tool usage”. A fixed pipeline that calls tools is therefore not an agent to Anthropic. NIST’s characterisation turns instead on autonomy and external state, and explicitly contemplates orchestration software driving multiple sub-agents.
Whether a language model is required. NIST requires “at least one generative AI model”. OWASP’s headline definition names no model type — “an intelligent software system designed to perceive its environment, reason about it, make decisions, and take actions to achieve specific objectives autonomously”. OWASP states that in its own voice, and then offers Russell and Norvig’s definition separately, as a more specific one; neither names a language model. The same document records that the OpenAI Gym “helped drive the first wave of Agentic AI” before large language models “revolutionized” it, and says modern agents use LLMs as their reasoning engines to decide control flow. So OWASP holds both the wide definition and the narrow practice.
Autonomy and planning recur across these sources, but as components rather than tests. Memory is thinner than that. Of the seven sources this page reads for a definition, three mention agent memory at all, and only OWASP lists it as a capability — behind a hedge, “There are many ways to describe an agent, but typically, an agent or agentic AI system will exhibit the following elements”. Neither OpenAI document mentions memory once; NIST’s single mention is memory safety in ordinary software, and Chan’s is a cognitive-science conception the paper sets aside. No source in this set makes memory a condition of membership.
Questions and answers
What is the difference between an AI agent and a chatbot?
An AI agent takes actions that change something outside itself; a chatbot returns text to a person who then decides what to do. NIST's Center for AI Standards and Innovation scopes its request for information to the security of AI agent systems "capable of taking actions that affect external state, i.e., persistent changes outside of the AI agent system itself", and sets aside chatbots and retrieval-augmented generation systems "that are not orchestrated to act autonomously". The qualifier matters: a chat interface does not disqualify a system, and the same assistant becomes an agent once it is wired to tools that act.
Does an AI agent have to be built on a language model?
Sources disagree. NIST's characterisation requires "at least one generative AI model and scaffolding software that equips the model with tools". OWASP's Agentic Security Initiative defines an agent without naming any model type — "an intelligent software system designed to perceive its environment, reason about it, make decisions, and take actions to achieve specific objectives autonomously" — and separately cites Russell and Norvig's definition, which names no model type either. It records that the OpenAI Gym helped drive the first wave of agentic AI before large language models "revolutionized" it. In current security writing the term almost always means a system built on a language model.
Is a fixed multi-step pipeline that calls tools an AI agent?
Not to Anthropic, which separates workflows from agents. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths"; agents are "systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks". NIST does not draw that line. Its characterisation turns on autonomous action affecting external state, and expressly allows agent systems "containing multiple sub-agents with software that orchestrates their interactions". Whether a given pipeline counts therefore depends on whose definition a claim is using.
Sources
- Request for Information Regarding Security Considerations for Artificial Intelligence AgentsCenter for AI Standards and Innovation, NIST, US Department of Commerce, 8 Jan 2026
- When Context Gets Root: Privilege Escalation in LLM HarnessesarXiv, 27 Aug 2026
- Re: Request for Information: Security Considerations for Artificial Intelligence Agents, Docket No. NIST-2025-0035Anthropic, 9 Mar 2026
- Building effective agentsAnthropic, 19 Dec 2024
- A practical guide to building agentsOpenAI
- Practices for Governing Agentic AI SystemsOpenAI
- Harms from Increasingly Agentic Algorithmic SystemsFAccT '23: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 12 Jun 2023
- Agentic AI - Threats and Mitigations, Version 1.1, December 2025OWASP Top 10 for LLM Apps & Gen AI, Agentic Security Initiative
Guides that use this term
Related topics
In the news
- Carriers and Amazon stop the Meta AI agent Muse
- The type of failure shows if a harness or weights repair an agent
- Thousands of AI agents in a test network moved into opposite camps
- Anthropic stops live internet access for its internal tests
- Claude can operate a maximum of 1,000 agents at the same time
- The Aave MCP server connects to the MetaMask Agent Wallet
- Meta and Sierra make an open protocol for personal AI agents
- Wikimedia reports that OpenAI agents sent millions of API requests
- In a study, AI agents increased prices for users with more wealth
- The AI tool REA has more than 10,400 stars on GitHub
- About 700 AI agents attacked Hugging Face, reports say
- A new step helps AI agents end a deadlock in 5 of 5 tests
- Tests show that the reward has a small effect on LLM agents
- Arbiter gives 152 results for the system prompts of 3 coding agents
- Models give no report of their errors in 67.1% of AI agent tests
- The LLM-IDEA agent tells when more experiments cannot help
- An OpenAI agent had unauthorised access to a Medicare portal
- Microsoft makes MXC available to limit the files an AI agent can use
- Incognia started a fraud check for AI agents
- Goodfire says its monitor for AI agents costs $51 for 1,500 sessions
- AI agents wrote GPU code with 1.57× the speed of `torch.compile`
- Crossmint made one toolkit for payments by AI agents
- One chat message gave Zenity control of all AgentCore agents
- The first AWS patch for AgentCore agents did not remove the risk
- President Lee thinks AI agents were used in attacks on 7 or more firms
- AI agents and humans get the same average, with different decisions
- A new attack agent goes around the guardrails of 4 AI models
- Agents in a multi-agent system start to have the same behavior
- An AI agent gives high confidence on 56% of high-risk tasks
- Reasoning AI agents do not select at random with the same inputs