Chuyển đến nội dung
Bản đồ AI

Tác tử tự chủ

Chia nhỏ mục tiêu và hành động đến khi xong

Tác nhân & gọi công cụTrung cấp #36
đầu vàoVăn bảnHành động

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NĂNG LỰC NÀY NGHĨA LÀ GÌ

Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.

Làm ra sao về mặt kỹ thuật

The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.

Sản phẩm tiêu biểu

5

Tổ chức liên quan

Cách dùng tiêu biểu

  • Fixing defects and submitting patches
  • Orchestrating and running data pipelines
  • Researching and assembling a report
  • Repetitive operations and scripted tasks

Đánh giá nó tốt hay không thế nào

Task completion rate
Share of problems solved end to end, as in SWE-bench-style evaluation
Steps and cost
Calls and compute spent to solve the task
Human takeover rate
Share requiring a human to step in mid-task

Ranh giới và điểm khó

  • On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
  • It can declare success without verifying: the task is announced as done while the check never ran
  • It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed

Các khái niệm đằng sau