自律エージェント
目標を分解し、完了まで連続して行動する
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
技術的にどう実現するか
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
代表的な製品
5Devin
2024シェル・エディタ・ブラウザを備えた自律型コーディングエージェント
Claude Code
2025ターミナルから多段のコーディング作業をこなすエージェント
LangChain
2022モデル・ツール・検索をつなぐ
Cursor
2023リポジトリ全体を文脈にするデスクトップコードエディタ
ChatGPT
2022大規模言語モデルを誰もが使える対話画面にした
関連する組織
代表的な用途
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
どう評価するか
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
限界と難しさ
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed