コンピュータ操作エージェント
人がするように画面を見て、クリックし、打ち込む
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a task to be done through a user interface — add this column to the sheet — and outputs a sequence of mouse and keyboard actions, with screenshots as the observation at each step. It needs no software API, so it can drive legacy systems that expose none. Unlike an autonomous agent its actions land on a graphical interface rather than code or APIs.
技術的にどう実現するか
The model treats screenshots as observations and emits actions such as click coordinates, scrolling and typing, advancing in a repeated screenshot–action loop. Training data comes from recordings of human operations and trajectories annotated on synthetic tasks, and evaluation runs on task sets over real websites and desktop apps. To reduce risk, deployments usually run in isolated virtual machines or constrained accounts with confirmation points on sensitive actions.
代表的な製品
5Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Devin
2024シェル・エディタ・ブラウザを備えた自律型コーディングエージェント
ChatGPT
2022大規模言語モデルを誰もが使える対話画面にした
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Cursor
2023リポジトリ全体を文脈にするデスクトップコードエディタ
関連する組織
代表的な用途
- Testing and accepting web flows
- Data entry into legacy systems without APIs
- Automating repetitive back-office operations
- Operating interfaces on behalf of users with accessibility needs
限界と難しさ
- Clicks based on screenshot pixel coordinates miss as soon as resolution or page zoom changes
- Captchas, pop-ups, long scrolling and lazy loading make it stall or repeat ineffective actions
- With no reliable intermediate feedback, a wrong click corrupts real accounts and data, so isolated execution is mandatory