Autonome Agenten
Ein Ziel zerlegen und bis zur Erledigung handeln
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
Wie sie technisch umgesetzt wird
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Repräsentative Produkte
5Devin
2024Ein autonomer Coding-Agent mit eigener Shell, Editor und Browser
Claude Code
2025Ein Coding-Agent, der mehrstufige Aufgaben im Terminal erledigt
LangChain
2022Modelle, Werkzeuge und Retrieval verketten
Cursor
2023Ein Desktop-Code-Editor, der das ganze Repository als Kontext nutzt
ChatGPT
2022Das Chatfenster, das ein großes Sprachmodell für alle zugänglich machte
Beteiligte Organisationen
Typische Verwendungen
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
Wie sie bewertet wird
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Grenzen und schwierige Punkte
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Konzepte dahinter
Agenten und Werkzeugnutzung
Ein Modell soll nicht nur antworten, sondern suchen, APIs aufrufen, Code ausführen – und anhand des Ergebnisses den nächsten Schritt wählen
Bestärkendes Lernen aus menschlichem Feedback
Wenn die „gute Antwort" keine Formel hat, übernehmen Menschen die Rolle der Belohnungsfunktion
Sicherheit, Alignment und Prompt-Injection
Ein Modell optimiert den Proxy in der Verlustfunktion, nie das eigentlich Gewollte – die Lücke dazwischen ist das ganze Alignment-Problem