Autonomous Agents
Break down a goal and keep acting until it is done
WHAT THIS CAPABILITY MEANS
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
How it is done
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Representative products
5Devin
2024An autonomous coding agent with its own shell, editor and browser
Claude Code
2025A coding agent that carries out multi-step tasks from the terminal
LangChain
2022Chain models, tools and retrieval together
Cursor
2023A desktop code editor that treats the whole repository as context
ChatGPT
2022The chat window that put a large language model in everyone’s hands
Organizations involved
Typical uses
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
How it is evaluated
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Limits and hard parts
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Concepts behind it
Agents & Tool Use
Let a model do more than answer: search, call APIs, run code — and decide the next step from what came back
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Safety, Alignment & Prompt Injection
A model optimises the proxy we wrote into the loss, never the thing we actually want — the gap between them is the whole alignment problem