Tác tử tự chủ
Chia nhỏ mục tiêu và hành động đến khi xong
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
Làm ra sao về mặt kỹ thuật
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Sản phẩm tiêu biểu
5Devin
2024Tác nhân lập trình tự chủ có shell, trình soạn thảo và trình duyệt riêng
Claude Code
2025Tác nhân lập trình hoàn thành việc nhiều bước ngay trong terminal
LangChain
2022Ghép mô hình, công cụ và truy hồi lại với nhau
Cursor
2023Trình soạn thảo mã trên máy tính lấy cả kho mã làm ngữ cảnh
ChatGPT
2022Cửa sổ trò chuyện đưa mô hình ngôn ngữ lớn đến tay mọi người
Tổ chức liên quan
Cách dùng tiêu biểu
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
Đánh giá nó tốt hay không thế nào
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Ranh giới và điểm khó
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Các khái niệm đằng sau
Tác tử và sử dụng công cụ
Để mô hình không chỉ trả lời, mà còn tra cứu, gọi API, chạy mã — và quyết định bước tiếp theo từ kết quả nhận được
Học tăng cường từ phản hồi của con người
Khi «câu trả lời tốt» không viết được thành công thức, hãy để con người đóng vai hàm thưởng
An toàn, căn chỉnh và chèn lệnh (prompt injection)
Mô hình tối ưu chỉ số đại diện ghi trong hàm mất mát, không phải điều ta thực sự muốn — khoảng cách đó chính là toàn bộ vấn đề căn chỉnh