Definition · AI agents
Web agent
A web agent is an AI agent scoped to a browser: it perceives a web page — through its accessibility tree, DOM elements, or a screenshot — and acts on it by clicking, typing, navigating and submitting, toward a goal. That scope distinguishes it from a computer-use agent, the broader category that also drives desktop apps, files and a terminal.
Last reviewed
Key points
- A web agent is an AI agent scoped to a browser: it reads a page's structure or a screenshot, and acts by clicking, typing, navigating and submitting.
- Anthropic's own tool docs draw the scope line: browser use for a task that stays inside webpages, computer use for a task that needs a whole desktop.
- Asteroid AI states the same split — a browser agent runs only inside a browser; a computer-use agent can also drive desktop apps, files and a terminal.
- TinyFish defines a web agent as a system that perceives, reasons about and acts on web pages, adapting to interface changes rather than breaking like a scraper.
- The same scope that makes a web agent useful is the risk surface cross-site prompt injection and agent hijacking exploit — a framing none of the three vendor guides use themselves.
A web agent perceives a web page and acts on it — nothing more. That scope separates it from a fixed script, which reads a selector rather than the page, and from a computer-use agent, which reaches beyond the browser.
How it works
Anthropic’s browser use tool documentation describes the loop plainly: it “lets Claude navigate, read, and interact with webpages in a browser that your application runs,” working “through its structure (the accessibility tree, elements, forms, and tabs) and through screenshots and viewport coordinates.” Acting means the ordinary browser verbs — click, type, navigate, scroll, submit, switch a tab. TinyFish describes the same loop from the product side: a web agent “uses large language models to perceive, reason about, and take action on live web pages,” adapting when interfaces change rather than breaking like a scripted scraper.
A lab and a startup draw the line the same way. Anthropic: “choose browser use when the task stays inside webpages,” reaching for computer use only “when a task needs a whole desktop.” Asteroid AI: “A browser agent runs only inside a browser; a computer-use agent can also drive desktop apps, the file system, and the terminal.”
Why it matters
The same scope that makes a web agent useful is what cross-site prompt injection exploits. An attacker plants an instruction on one page; unable to tell instruction from data, the agent carries it out elsewhere in the same session — an agent hijacking. A web agent with account or payment actions available is what makes that instruction consequential.
None of the three vendor guides defining the term frame it this way. LayerX comes closest, warning that “a compromised agent could be used to exfiltrate sensitive data, hijack user sessions, or perform unauthorized actions” — but its scenario is a hijacked extension, not ordinary content turned into an instruction.
Questions and answers
What is the difference between a web agent and a computer-use agent?
Scope. A web agent works only inside a browser — reading a page's structure or a screenshot, then clicking, typing, navigating and submitting. A computer-use agent is the broader category: the same loop, but able to drive desktop applications, the file system and a terminal too. Asteroid AI states the line directly: "A browser agent runs only inside a browser; a computer-use agent can also drive desktop apps, the file system, and the terminal." Anthropic's own tooling draws the same line — use browser use "when the task stays inside webpages," and computer use "when a task needs a whole desktop."
How does a web agent "see" a web page?
Two ways, often both. It can read the page's structure — an accessibility tree of labelled elements it can act on directly, without first locating them in an image — or a screenshot, when the visual layout matters or the structure alone falls short. Anthropic's browser use tool documents both channels, describing Claude as working "with the page both through its structure (the accessibility tree, elements, forms, and tabs) and through screenshots and viewport coordinates."
Why does a web agent's scope matter for security?
Because everything it can read becomes something it might act on. A web agent that can click, submit and message doesn't just automate a task — it gives an attacker a lever: plant an instruction on any page it reads, and the agent, unable to tell instruction from data, may carry it out as an action. That is the mechanism cross-site prompt injection exploits. LayerX's own guide warns that a compromised agent "could be used to exfiltrate sensitive data, hijack user sessions, or perform unauthorized actions" — but its scenario is a hijacked browser extension, not an instruction hidden in ordinary content the agent was already reading.
Sources
- Browser use toolAnthropic
- What Are Browser Agents? A 2026 GuideAsteroid AI (Edward Upton, Founding Engineer), 10 May 2026
- What Is a Web Agent? The Complete Guide to AI Browser Agents in 2026TinyFish, 10 Apr 2026
- What Are AI Browser Agents and How to Build ThemLayerX Security (Or Eshed), 11 Nov 2025