推論と思考の連鎖
難問を中間ステップに分解して解く
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a problem requiring several steps — a maths problem, a logic puzzle, a task needing a plan — and returns the answer together with intermediate steps. Unlike free generation the goal is correctness rather than fluency, and unlike ordinary conversation it spends much of its compute on an internal scratchpad, showing the user mainly the result and a short rationale.
技術的にどう実現するか
One family relies on prompting: examples or instructions ask the model to think before answering, writing steps out explicitly — chain-of-thought. Another relies on training: large-scale reinforcement learning on verifiable problems such as maths and code teaches the model to generate a longer deliberation before committing. Self-consistency voting over several sampled chains, and combining deliberation with tool calls or search, are now common too.
代表的な製品
23o3
2025答える前に長い推論を重ね、推論時の計算で正答率を上げる
DeepSeek-R1
2025強化学習で推論の連鎖を訓練し、MIT ライセンスで重みを公開した推論モデル
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
MiniMax-M
2025ハイブリッド注意と100万トークン文脈をもつオープンウェイトの推論モデル
DeepSeek-V3
2024総パラメータ 671B、1 トークンあたり 37B だけを活性化するオープンウェイト MoE
GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Jamba
2024状態空間モデルと Transformer を混合したオープンウェイトモデル
DBRX
2024DatabricksのオープンウェイトMoE言語モデル
Mistral Large
2024欧州のオープンウェイト研究所による商用フラッグシップモデル
Gemini
2024検索・オフィス・マルチモーダルモデルを一つの対話入口に集約
Grok
2023ソーシャルプラットフォームのデータと結びついた対話モデル
Yi
2023中英バイリンガルで超長文脈版もあるオープンウェイトモデル
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
Doubao
2023バイトダンスの汎用対話モデルとアプリ
Phi
2023小さな規模と厳選データで作られた高効率な小モデル
Step
2023マルチモーダルとオンデバイスを狙う汎用モデル群
Baichuan
2023中国語シーンに向けたオープンウェイトの汎用モデル
GLM
2023自己回帰空白穴埋め事前学習から始まった中国語の汎用モデル
Llama
2023オープンウェイト路線を主流にした汎用モデル群
ChatGPT
2022大規模言語モデルを誰もが使える対話画面にした
Pangu
2020華為の盤古シリーズ基盤モデル
関連する組織
代表的な用途
- Solving maths and physics problems
- Multi-step planning and scheduling
- Code and data problems needing derivation
- Reasoned choices under constraints
どう評価するか
- Math benchmark accuracy
- Correct-answer rate on sets such as GSM8K, MATH, AIME
- pass@k
- Share of problems solved by at least one of k samples
- Chain-of-thought faithfulness
- Whether the written steps actually reflect how the answer was reached
限界と難しさ
- The written reasoning can be unfaithful: the answer comes first and a plausible rationale is added after
- In long chains an early misstep propagates, yet the final answer still sounds assured
- Deliberating at length on trivial questions inflates latency and cost