Agentes autônomos
Decompor um objetivo e agir até concluí-lo
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes a higher-level goal — fix this issue — and outputs a sequence of actions, adjusting the next step from feedback along the way. Unlike tool use it orchestrates many calls into a plan, deciding the order and when to stop; unlike a computer-use agent it works mainly through code and APIs rather than a screen interface.
Como é feita tecnicamente
The typical structure is a think–act–observe loop: the model plans a step, calls a tool or runs code, reads the result and revises the plan, repeating until it judges the task done. Capability comes from three things stacked together: a long-context language model, immediate feedback from an executable environment (compilation, tests, errors), and outer controls for safety and budget such as sandboxes, step limits and human checkpoints.
Produtos representativos
5Devin
2024Um agente de código autônomo com shell, editor e navegador próprios
Claude Code
2025Um agente de código que executa tarefas de várias etapas no terminal
LangChain
2022Encadeie modelos, ferramentas e recuperação
Cursor
2023Um editor de código de desktop que usa o repositório inteiro como contexto
ChatGPT
2022A janela de chat que colocou um grande modelo de linguagem nas mãos de todos
Organizações relacionadas
Usos típicos
- Fixing defects and submitting patches
- Orchestrating and running data pipelines
- Researching and assembling a report
- Repetitive operations and scripted tasks
Como avaliar se funciona bem
- Task completion rate
- Share of problems solved end to end, as in SWE-bench-style evaluation
- Steps and cost
- Calls and compute spent to solve the task
- Human takeover rate
- Share requiring a human to step in mid-task
Limites e dificuldades
- On long tasks errors accumulate, and after drifting off course the agent rarely recovers on its own
- It can declare success without verifying: the task is announced as done while the check never ran
- It lacks a reliable internal brake for high-risk actions such as deleting, deploying or paying, so external guardrails are needed
Conceitos por trás
Agentes e uso de ferramentas
Que um modelo não só responda: pesquise, chame APIs, execute código — e decida o próximo passo a partir do resultado
Aprendizado por reforço a partir de feedback humano
Quando a «boa resposta» não cabe numa fórmula, deixe as pessoas servirem de recompensa
Segurança, alinhamento e injeção de prompt
O modelo otimiza o proxy escrito na perda, nunca o que realmente queremos — essa lacuna é todo o problema de alinhamento