컴퓨터 사용 에이전트
사람처럼 화면을 보고 클릭하고 입력한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a task to be done through a user interface — add this column to the sheet — and outputs a sequence of mouse and keyboard actions, with screenshots as the observation at each step. It needs no software API, so it can drive legacy systems that expose none. Unlike an autonomous agent its actions land on a graphical interface rather than code or APIs.
기술적으로 구현하는 방법
The model treats screenshots as observations and emits actions such as click coordinates, scrolling and typing, advancing in a repeated screenshot–action loop. Training data comes from recordings of human operations and trajectories annotated on synthetic tasks, and evaluation runs on task sets over real websites and desktop apps. To reduce risk, deployments usually run in isolated virtual machines or constrained accounts with confirmation points on sensitive actions.
대표 제품
5Claude
2023긴 문맥과 안전 정렬로 알려진 범용 대화 모델
Devin
2024셸·편집기·브라우저를 갖춘 자율 코딩 에이전트
ChatGPT
2022대규모 언어 모델을 누구나 쓰는 대화 창으로 만들었다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Cursor
2023저장소 전체를 문맥으로 삼는 데스크톱 코드 편집기
관련 기관
대표적 용도
- Testing and accepting web flows
- Data entry into legacy systems without APIs
- Automating repetitive back-office operations
- Operating interfaces on behalf of users with accessibility needs
경계와 난점
- Clicks based on screenshot pixel coordinates miss as soon as resolution or page zoom changes
- Captchas, pop-ups, long scrolling and lazy loading make it stall or repeat ineffective actions
- With no reliable intermediate feedback, a wrong click corrupts real accounts and data, so isolated execution is mandatory