本文へスキップ
AI図鑑

推論と思考の連鎖

難問を中間ステップに分解して解く

言語と知識中級 #08
入力テキストテキスト

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

この能力とは何か

Takes a problem requiring several steps — a maths problem, a logic puzzle, a task needing a plan — and returns the answer together with intermediate steps. Unlike free generation the goal is correctness rather than fluency, and unlike ordinary conversation it spends much of its compute on an internal scratchpad, showing the user mainly the result and a short rationale.

技術的にどう実現するか

One family relies on prompting: examples or instructions ask the model to think before answering, writing steps out explicitly — chain-of-thought. Another relies on training: large-scale reinforcement learning on verifiable problems such as maths and code teaches the model to generate a longer deliberation before committing. Self-consistency voting over several sampled chains, and combining deliberation with tool calls or search, are now common too.

代表的な製品

23

o3

2025
OpenAI

答える前に長い推論を重ね、推論時の計算で正答率を上げる

モデル クローズド
テキスト画像テキスト

DeepSeek-R1

2025
DeepSeek

強化学習で推論の連鎖を訓練し、MIT ライセンスで重みを公開した推論モデル

モデル オープンウェイト
テキストテキスト

Gemini

2023
Google DeepMind

ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル

モデル クローズド
テキスト画像音声動画テキスト

Claude

2023
Anthropic

長い文脈と安全性の調整で知られる汎用対話モデル

モデル クローズド
テキスト画像テキスト

Qwen

2023
Alibaba (Qwen)

多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群

モデル オープンウェイト
テキスト画像テキスト

MiniMax-M

2025
MiniMax

ハイブリッド注意と100万トークン文脈をもつオープンウェイトの推論モデル

モデル オープンウェイト
テキストテキスト

DeepSeek-V3

2024
DeepSeek

総パラメータ 671B、1 トークンあたり 37B だけを活性化するオープンウェイト MoE

モデル オープンウェイト
テキストテキスト

GPT-4o

2024
OpenAI

ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う

モデル クローズド
テキスト画像音声テキスト音声

Jamba

2024
AI21 Labs

状態空間モデルと Transformer を混合したオープンウェイトモデル

モデル オープンウェイト
テキストテキスト

DBRX

2024
Databricks

DatabricksのオープンウェイトMoE言語モデル

モデル オープンウェイト
テキストテキスト

Mistral Large

2024
Mistral AI

欧州のオープンウェイト研究所による商用フラッグシップモデル

モデル クローズド
テキストテキスト

Gemini

2024
Google DeepMind

検索・オフィス・マルチモーダルモデルを一つの対話入口に集約

アプリ フリーミアム
テキスト画像音声テキスト画像音声

Grok

2023
xAI

ソーシャルプラットフォームのデータと結びついた対話モデル

モデル クローズド
テキスト画像テキスト

Yi

2023
01.AI

中英バイリンガルで超長文脈版もあるオープンウェイトモデル

モデル オープンウェイト
テキストテキスト

Hunyuan

2023
Tencent (Hunyuan)

開放ウェイト版を含むテンセントの汎用モデル群

モデル オープンウェイト
テキスト画像テキスト

Doubao

2023
ByteDance (Seed)

バイトダンスの汎用対話モデルとアプリ

モデル クローズド
テキスト画像テキスト

Phi

2023
Microsoft

小さな規模と厳選データで作られた高効率な小モデル

モデル オープンウェイト
テキストテキスト

Step

2023
StepFun

マルチモーダルとオンデバイスを狙う汎用モデル群

モデル クローズド
テキスト画像テキスト

Baichuan

2023
Baichuan AI

中国語シーンに向けたオープンウェイトの汎用モデル

モデル オープンウェイト
テキストテキスト

GLM

2023
Zhipu AI

自己回帰空白穴埋め事前学習から始まった中国語の汎用モデル

モデル クローズド
テキストテキスト

Llama

2023
Meta AI (FAIR)

オープンウェイト路線を主流にした汎用モデル群

モデル オープンウェイト
テキストテキスト

ChatGPT

2022
OpenAI

大規模言語モデルを誰もが使える対話画面にした

アプリ フリーミアム
テキスト画像音声テキスト画像音声

Pangu

2020
Huawei

華為の盤古シリーズ基盤モデル

モデル クローズド
テキストテキスト

関連する組織

代表的な用途

  • Solving maths and physics problems
  • Multi-step planning and scheduling
  • Code and data problems needing derivation
  • Reasoned choices under constraints

どう評価するか

Math benchmark accuracy
Correct-answer rate on sets such as GSM8K, MATH, AIME
pass@k
Share of problems solved by at least one of k samples
Chain-of-thought faithfulness
Whether the written steps actually reflect how the answer was reached

限界と難しさ

  • The written reasoning can be unfaithful: the answer comes first and a plausible rationale is added after
  • In long chains an early misstep propagates, yet the final answer still sounds assured
  • Deliberating at length on trivial questions inflates latency and cost

背景にある概念