Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
WHAT THIS CAPABILITY MEANS
Takes a problem requiring several steps — a maths problem, a logic puzzle, a task needing a plan — and returns the answer together with intermediate steps. Unlike free generation the goal is correctness rather than fluency, and unlike ordinary conversation it spends much of its compute on an internal scratchpad, showing the user mainly the result and a short rationale.
How it is done
One family relies on prompting: examples or instructions ask the model to think before answering, writing steps out explicitly — chain-of-thought. Another relies on training: large-scale reinforcement learning on verifiable problems such as maths and code teaches the model to generate a longer deliberation before committing. Self-consistency voting over several sampled chains, and combining deliberation with tool calls or search, are now common too.
Representative products
23o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
DeepSeek-R1
2025A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
Gemini
2023A natively multimodal general model built for very long context
Claude
2023A general chat model known for long context and safety alignment
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Jamba
2024An open-weight model that mixes a state-space model with a Transformer
DBRX
2024Databricks’ open-weights MoE language model
Mistral Large
2024The flagship commercial model from a European open-weights lab
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
Grok
2023A chat model tied to a social platform’s live data
Yi
2023A bilingual Chinese–English open-weight model with very-long-context versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
Doubao
2023ByteDance’s general chat model and application
Phi
2023Small, efficient models built from curated data
Step
2023A general model family aimed at multimodality and on-device use
Baichuan
2023An open-weight general model aimed at Chinese-language use
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Llama
2023The model family that made the open-weight route mainstream
ChatGPT
2022The chat window that put a large language model in everyone’s hands
Pangu
2020Huawei’s Pangu family of foundation models
Organizations involved
Typical uses
- Solving maths and physics problems
- Multi-step planning and scheduling
- Code and data problems needing derivation
- Reasoned choices under constraints
How it is evaluated
- Math benchmark accuracy
- Correct-answer rate on sets such as GSM8K, MATH, AIME
- pass@k
- Share of problems solved by at least one of k samples
- Chain-of-thought faithfulness
- Whether the written steps actually reflect how the answer was reached
Limits and hard parts
- The written reasoning can be unfaithful: the answer comes first and a plausible rationale is added after
- In long chains an early misstep propagates, yet the final answer still sounds assured
- Deliberating at length on trivial questions inflates latency and cost
Concepts behind it
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away