構造化出力
決められた schema どおりの JSON を出させる
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes natural language plus a target structure — field names, types, required flags — and outputs text that strictly conforms, such as JSON, XML or a table. Unlike information extraction the structure is supplied by the caller rather than inferred from the text; unlike tool use it produces the data itself, not a request to invoke some interface.
技術的にどう実現するか
The basic approach puts the schema and examples in the prompt and asks the model to fill them; more reliable is constrained decoding, which at each step allows only tokens that keep the structure valid, ruling out syntax errors during generation. Fine-tuning on many text-to-structure pairs familiarises the model with common field naming and type conventions. Large schemas can be split across several calls or filled in layers.
代表的な製品
5GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
関連する組織
代表的な用途
- Converting natural language into API payloads
- Auto-filling forms and order fields
- Data cleaning and normalisation pipelines
- Passing structured messages between agents
どう評価するか
- Schema validity rate
- Share of outputs that parse and satisfy the schema
- Field exact match
- Share of field values matching the reference
- Value-level F1
- Scores elements in lists and nested structures
限界と難しさ
- With many fields or long enums it drops fields, picks the wrong value or invents new ones
- Deep nesting and complex types such as unions or nullable arrays produce type errors
- Tight format constraints squeeze reasoning, and the same model reasons less well under them