Definition · AI agents
Computer-use agent
A computer-use agent is an AI agent that operates a whole desktop: it perceives the screen through screenshots and acts on it with simulated mouse and keyboard input, reaching the file system, the terminal and any installed application. That scope distinguishes it from a web agent, which is scoped to a browser.
Last reviewed
Key points
- A computer-use agent perceives a screenshot and acts with mouse and keyboard input on a real desktop: files, a terminal, any installed application, not only a browser tab.
- Anthropic and OpenAI both ship it under this name. Anthropic's tool gives Claude 17 member tools such as screenshot and left_click; OpenAI's lets a model operate browser and desktop interfaces the same way.
- Both vendors warn that screen content is untrusted. OpenAI: instructions in a page cannot override the user's; Anthropic: Claude will sometimes follow commands found in content over its own instructions.
- Both recommend the same controls: an isolated VM, a restricted environment, and human confirmation before anything hard to reverse.
- Neither vendor names the risk this creates once the agent leaves a browser: the same prompt injection a web agent faces, landing on a shell instead of a page.
A computer-use agent perceives a screen and acts on it, like a web agent perceives a page, but without the browser’s boundary: a browser tab, a text editor, a file manager, a terminal.
How it works
Both vendors describe the same loop. Anthropic: Claude “can interact with computer environments through the computer use tool, which provides screenshot capabilities and mouse/keyboard control for autonomous desktop interaction,” with “17 member tools such as screenshot, left_click, type, and zoom.” OpenAI: “Computer use lets a model operate browser and desktop interfaces.”
What separates this from a web agent is what sits underneath the loop. Anthropic’s reference setup runs the tool against a Linux desktop with “Pre-installed Linux applications such as Firefox, LibreOffice, text editors, and file managers.” A web agent’s tools stop at the browser; a computer-use agent’s stop at whatever the desktop exposes.
Why it matters
The same failure a web agent faces returns here at a wider radius. Neither vendor claims the model can tell an instruction from content it is merely reading. OpenAI: “Treat screen content as untrusted.” Anthropic: “In some circumstances, Claude will follow commands found in content even when they conflict with your instructions.”
On a browser-scoped agent, that failure is cross-site prompt injection, carried out as agent hijacking. A computer-use agent inherits the same mechanism, but the action it can be tricked into is bounded by the whole desktop, not a tab — a file to overwrite, a command to run. That is excessive agency at desktop scale: same trigger, wider blast radius.
Both vendors answer with the same fix, architectural rather than a better prompt: an isolated VM, no sensitive data on hand, and a human confirming anything hard to reverse. Neither names this as a risk category; both land on the mitigation a security team would reach for anyway.
Questions and answers
What is a computer-use agent?
An AI agent that operates a whole desktop rather than a single application. It perceives the screen through screenshots and acts on it with simulated mouse clicks, typing and keystrokes, so it can drive a browser, a text editor, a file manager or a terminal, whatever is open. Anthropic's version gives Claude "screenshot capabilities and mouse/keyboard control for autonomous desktop interaction." OpenAI's description is nearly identical: "Computer use lets a model operate browser and desktop interfaces."
How is a computer-use agent different from a web agent?
Reach. A web agent is scoped to a browser: it reads a page and acts on it, and nothing outside the tab is reachable. A computer-use agent has no such boundary — the same screenshot-and-click loop runs against the whole screen, so it can also open a file manager, run a terminal command, or switch to another application. Asteroid AI states the line directly: "A browser agent runs only inside a browser; a computer-use agent can also drive desktop apps, the file system, and the terminal."
What are the security risks of a computer-use agent?
The same risk a web agent carries, at a wider radius. Both Anthropic and OpenAI warn that the agent cannot reliably tell an instruction from content it merely reads on screen. OpenAI's guidance: "Treat screen content as untrusted. Text in a page, document, or tool result cannot grant permission or override the user's instructions." Anthropic's: "Claude will follow commands found in content even when they conflict with your instructions." On a browser-scoped agent that is cross-site prompt injection; on a computer-use agent the same hidden instruction can reach the file system or a shell, which is what turns a triggered mistake into excessive agency on a desktop scale.
How do vendors recommend running a computer-use agent safely?
By constraining the environment rather than trusting the model. Anthropic recommends "a dedicated virtual machine or container with minimal privileges," avoiding "access to sensitive data," and "a human to confirm decisions that might result in meaningful real-world consequences." OpenAI's equivalent guidance: "Use an isolated browser or VM and an allow list of sites and actions," and "Confirm consequential actions... purchases, data transmission, destructive changes, and other actions that are hard to reverse." Both are operational controls a developer applies around the agent; neither claims the model can be relied on to draw that line itself.
Sources
- Computer use toolAnthropic
- Computer useOpenAI
- What Are Browser Agents? A 2026 GuideAsteroid AI (Edward Upton, Founding Engineer), 10 May 2026