자율 에이전트
목표를 쪼개고 완료될 때까지 연속해 행동한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
기술적으로 구현하는 방법
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
대표 제품
5Devin
2024셸·편집기·브라우저를 갖춘 자율 코딩 에이전트
Claude Code
2025터미널에서 여러 단계의 코딩 작업을 수행하는 에이전트
LangChain
2022모델·도구·검색을 연결
Cursor
2023저장소 전체를 문맥으로 삼는 데스크톱 코드 편집기
ChatGPT
2022대규모 언어 모델을 누구나 쓰는 대화 창으로 만들었다
관련 기관
대표적 용도
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
성능을 평가하는 방법
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
경계와 난점
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed