Computer-Nutzungs-Agenten
Wie ein Mensch auf den Bildschirm sehen, klicken und tippen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a task to be done through a user interface — add this column to the sheet — and outputs a sequence of mouse and keyboard actions, with screenshots as the observation at each step. It needs no software API, so it can drive legacy systems that expose none. Unlike an autonomous agent its actions land on a graphical interface rather than code or APIs.
Wie sie technisch umgesetzt wird
The model treats screenshots as observations and emits actions such as click coordinates, scrolling and typing, advancing in a repeated screenshot–action loop. Training data comes from recordings of human operations and trajectories annotated on synthetic tasks, and evaluation runs on task sets over real websites and desktop apps. To reduce risk, deployments usually run in isolated virtual machines or constrained accounts with confirmation points on sensitive actions.
Repräsentative Produkte
5Claude
2023Ein allgemeines Dialogmodell, bekannt für langen Kontext und Sicherheitsausrichtung
Devin
2024Ein autonomer Coding-Agent mit eigener Shell, Editor und Browser
ChatGPT
2022Das Chatfenster, das ein großes Sprachmodell für alle zugänglich machte
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Cursor
2023Ein Desktop-Code-Editor, der das ganze Repository als Kontext nutzt
Beteiligte Organisationen
Typische Verwendungen
- Testing and accepting web flows
- Data entry into legacy systems without APIs
- Automating repetitive back-office operations
- Operating interfaces on behalf of users with accessibility needs
Grenzen und schwierige Punkte
- Clicks based on screenshot pixel coordinates miss as soon as resolution or page zoom changes
- Captchas, pop-ups, long scrolling and lazy loading make it stall or repeat ineffective actions
- With no reliable intermediate feedback, a wrong click corrupts real accounts and data, so isolated execution is mandatory
Konzepte dahinter
Agenten und Werkzeugnutzung
Ein Modell soll nicht nur antworten, sondern suchen, APIs aufrufen, Code ausführen – und anhand des Ergebnisses den nächsten Schritt wählen
Prompt-Engineering und Alignment
Ein Modell hilfreich, ehrlich und harmlos zu machen ist schwieriger, als es einfach größer zu machen
Sicherheit, Alignment und Prompt-Injection
Ein Modell optimiert den Proxy in der Verlustfunktion, nie das eigentlich Gewollte – die Lücke dazwischen ist das ganze Alignment-Problem