Agentes autónomos
Descomponer un objetivo y actuar hasta completarlo
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
Cómo se consigue técnicamente
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Productos representativos
5Devin
2024Un agente de código autónomo con su propio shell, editor y navegador
Claude Code
2025Un agente de código que resuelve tareas de varios pasos desde la terminal
LangChain
2022Encadena modelos, herramientas y recuperación
Cursor
2023Un editor de código de escritorio que usa todo el repositorio como contexto
ChatGPT
2022La ventana de chat que puso un gran modelo de lenguaje en manos de todos
Organizaciones relacionadas
Usos típicos
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
Cómo se evalúa
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Límites y dificultades
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Conceptos detrás
Agentes y uso de herramientas
Que un modelo no solo responda: consulte, llame APIs, ejecute código — y decida el siguiente paso según el resultado
Aprendizaje por refuerzo a partir de retroalimentación humana
Cuando la «buena respuesta» no cabe en una fórmula, deja que las personas actúen como recompensa
Seguridad, alineación e inyección de prompts
El modelo optimiza el proxy escrito en la pérdida, nunca lo que de verdad queremos: esa brecha es todo el problema de alineación