Автономные агенты
Разбить цель и действовать до её выполнения
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
Как это устроено
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Примеры продуктов
5Devin
2024Автономный агент-программист со своей оболочкой, редактором и браузером
Claude Code
2025Агент-программист, выполняющий многошаговые задачи прямо из терминала
LangChain
2022Связывание моделей, инструментов и поиска
Cursor
2023Настольный редактор кода, использующий весь репозиторий как контекст
ChatGPT
2022Окно диалога, которое сделало большую языковую модель доступной каждому
Связанные организации
Типичное применение
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
Как её оценивают
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Границы и трудности
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Концепции в основе
Агенты и использование инструментов
Пусть модель не только отвечает, но и ищет, вызывает API, запускает код — и выбирает следующий шаг по полученному результату
Обучение с подкреплением на основе обратной связи от людей
Когда «хороший ответ» не выразить формулой, роль функции награды берут на себя люди
Безопасность, выравнивание и инъекция промптов
Модель оптимизирует прокси в функции потерь, а не то, что мы действительно хотим; зазор между ними и есть вся проблема выравнивания