मुख्य सामग्री पर जाएँ

स्वायत्त एजेंट

लक्ष्य को बाँटकर पूरा होने तक कार्य करना

एजेंट और टूल उपयोगमध्यवर्ती #36
इनपुटटेक्स्टएक्शन

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.

तकनीकी रूप से कैसे

The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.

प्रतिनिधि उत्पाद

5

संबंधित संस्थान

सामान्य उपयोग

  • Fixing defects and submitting patches
  • Orchestrating and running data pipelines
  • Researching and assembling a report
  • Repetitive operations and scripted tasks

इसका मूल्यांकन कैसे होता है

Task completion rate
Share of problems solved end to end, as in SWE-bench-style evaluation
Steps and cost
Calls and compute spent to solve the task
Human takeover rate
Share requiring a human to step in mid-task

सीमाएँ और कठिनाइयाँ

  • On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
  • It can declare success without verifying: the task is announced as done while the check never ran
  • It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed

इसके पीछे की अवधारणाएँ