Computer-Use Agents
See the screen, click and type like a person
WHAT THIS CAPABILITY MEANS
Takes a task to be done through a user interface — add this column to the sheet — and outputs a sequence of mouse and keyboard actions, with screenshots as the observation at each step. It needs no software API, so it can drive legacy systems that expose none. Unlike an autonomous agent its actions land on a graphical interface rather than code or APIs.
How it is done
The model treats screenshots as observations and emits actions such as click coordinates, scrolling and typing, advancing in a repeated screenshot–action loop. Training data comes from recordings of human operations and trajectories annotated on synthetic tasks, and evaluation runs on task sets over real websites and desktop apps. To reduce risk, deployments usually run in isolated virtual machines or constrained accounts with confirmation points on sensitive actions.
Representative products
5Claude
2023A general chat model known for long context and safety alignment
Devin
2024An autonomous coding agent with its own shell, editor and browser
ChatGPT
2022The chat window that put a large language model in everyone’s hands
Gemini
2023A natively multimodal general model built for very long context
Cursor
2023A desktop code editor that treats the whole repository as context
Organizations involved
Typical uses
- Testing and accepting web flows
- Data entry into legacy systems without APIs
- Automating repetitive back-office operations
- Operating interfaces on behalf of users with accessibility needs
Limits and hard parts
- Clicks based on screenshot pixel coordinates miss as soon as resolution or page zoom changes
- Captchas, pop-ups, long scrolling and lazy loading make it stall or repeat ineffective actions
- With no reliable intermediate feedback, a wrong click corrupts real accounts and data, so isolated execution is mandatory
Concepts behind it
Agents & Tool Use
Let a model do more than answer: search, call APIs, run code — and decide the next step from what came back
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Safety, Alignment & Prompt Injection
A model optimises the proxy we wrote into the loss, never the thing we actually want — the gap between them is the whole alignment problem