エージェントとツール利用
モデルに「答える」だけでなく、調べ、API を呼び、コードを走らせる力を持たせ、その結果から次の一手を決めさせる
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
定義
An agent is a system driven by a large model that autonomously decomposes a goal into multiple steps, interacts with the world by calling tools (functions, APIs, code interpreters, browsers), and decides each next step from the feedback it receives. Tool use is the core mechanism: the model emits a structured call request, the host executes it, and the result is returned to the model.
直観的な理解
A model without tools is an adviser who can only answer from memory; an agent hands that adviser a phone, a calculator, a filing cabinet and a runner. The adviser decides what to do and in what order, the tools do it. Risk comes with the power: an adviser who can place calls and edit files does far more damage when it misjudges than one who merely says the wrong thing.
The agent’s core loop: observe, reason, select and run a tool, evaluate the result — repeating until the task is done or a limit trips
Distribution of common agent failure causes (illustrative magnitude): tool use and loop control together exceed half, showing that reliable tool interfaces and stopping conditions matter more than raw cleverness
- Tool arguments or format errors · 34%
- Looping / failing to terminate · 24%
- Lost context or memory · 20%
- Planning errors or goal drift · 17%
- Permission or safety blocks · 5%
仕組み
- 01
Tool schemas and structured output
Each tool has a description: name, purpose, argument structure and types. The model no longer emits free text but a structured call (often JSON) matching the contract. Vague descriptions lead to misuse or missed tools — a tool’s manual is itself part of prompt engineering.
- 02
The observe–reason–act loop
Paradigms such as ReAct cycle the model through: observe the current state, reason about the next step, choose an action, read the result, and return to observing. The loop ends when the model judges the task done, or when a step, time or budget limit trips.
- 03
Memory and state management
A long task generates far more history than the context window holds. Working memory keeps the current step’s intermediate results; long-term memory writes key conclusions to external storage for later retrieval. When to compress, drop or write back decides whether the agent can finish a long workflow at all.
- 04
Error recovery and safety boundaries
Tools fail, arguments are wrong, models loop. A robust agent retries, falls back, swaps tools and self-corrects, and gates irreversible actions (transfers, deletions, sending messages) behind human confirmation and permission checks — every increment of autonomy should be matched by an increment of guardrail.
応用場面
- Coding assistants: reading files, editing code, running tests and continuing from failures
- Research assistants: searching, reading, extracting and assembling sourced reports
- Support ticket handling: looking up orders, calling refund APIs and updating ticket status
- Multi-agent collaboration: planner, executor and reviewer dividing work and checking one another
よくある誤解
- More autonomy means a larger risk surface. An agent that can perform irreversible actions pays for misjudgement in the real world, not merely in a wrong sentence.
- Errors compound in long tasks. A small deviation per step can snowball into a completely wrong objective after dozens of steps, and the model often fails to notice it has drifted.
- The context window is a hard constraint. History beyond it is truncated or forgotten; without explicit compression and retrieval, an agent “forgets” its own earlier decisions mid-workflow.
重要用語
- Function calling
- The model emitting structured arguments to invoke an external function
- ReAct
- A prompting paradigm alternating reasoning and action
- Tool
- One external capability an agent may invoke
- Guardrail
- Checks and constraints bounding what an agent may do