Tool Use & Function Calling
Let the model pick an API and fill its arguments
WHAT THIS CAPABILITY MEANS
Takes a user request and a list of available tools — names, descriptions, argument schemas — and outputs a structured call such as get_weather with city set to Shanghai. The model executes nothing itself; it selects a tool and assembles arguments, while the actual call is made by external code whose result is fed back for further reasoning.
How it is done
Fine-tuning on many examples of when to call, which tool to use and what arguments look like teaches the model to emit a call when it needs external information or action; at inference the tool definitions go into the prompt and the output is constrained to a valid structure. With many tools, retrieval first narrows the candidates, and a multi-turn loop that feeds results back handles complex tasks.
Representative products
10LangChain
2022Chain models, tools and retrieval together
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Claude
2023A general chat model known for long context and safety alignment
Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context
Command R
2024A commercial model built for retrieval augmentation and tool use
Mistral Large
2024The flagship commercial model from a European open-weights lab
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Llama
2023The model family that made the open-weight route mainstream
Organizations involved
Typical uses
- Fetching live data such as weather, prices or stock
- Performing business actions such as orders and messages
- Calling calculators, search and code executors
- Chaining several APIs for a composite task
How it is evaluated
- Function-selection accuracy
- Share of requests where the right tool was chosen
- Exact argument match
- Share with arguments exactly matching the reference
- End-to-end task success
- Whether the task completes once tools actually run
Limits and hard parts
- Slightly complex argument shapes — nested objects, arrays, optional fields — are easily filled wrongly
- As the tool list grows, wrong selections rise sharply, and similar-purpose tools are the most confusable
- Recovery after a failed call is unreliable, and the model often loops on the error instead of adapting
Concepts behind it
Agents & Tool Use
Let a model do more than answer: search, call APIs, run code — and decide the next step from what came back
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away