Skip to content
AI Atlas

Reasoning & Chain-of-Thought

Break a hard problem into intermediate steps

Language & knowledgeIntermediate #08
inTextText

WHAT THIS CAPABILITY MEANS

Takes a problem requiring several steps — a maths problem, a logic puzzle, a task needing a plan — and returns the answer together with intermediate steps. Unlike free generation the goal is correctness rather than fluency, and unlike ordinary conversation it spends much of its compute on an internal scratchpad, showing the user mainly the result and a short rationale.

How it is done

One family relies on prompting: examples or instructions ask the model to think before answering, writing steps out explicitly — chain-of-thought. Another relies on training: large-scale reinforcement learning on verifiable problems such as maths and code teaches the model to generate a longer deliberation before committing. Self-consistency voting over several sampled chains, and combining deliberation with tool calls or search, are now common too.

Representative products

23

o3

2025
OpenAI

It reasons at length before answering, trading inference-time compute for steadier accuracy

Model Closed
TextImageText

DeepSeek-R1

2025
DeepSeek

A reasoning model trained with RL on chains of thought, its weights open under the MIT licence

Model Open weights
TextText

Gemini

2023
Google DeepMind

A natively multimodal general model built for very long context

Model Closed
TextImageAudioVideoText

Claude

2023
Anthropic

A general chat model known for long context and safety alignment

Model Closed
TextImageText

Qwen

2023
Alibaba (Qwen)

An open-weight family spanning many sizes, with multimodal versions

Model Open weights
TextImageText

MiniMax-M

2025
MiniMax

An open-weight reasoning model with hybrid attention and a million-token context

Model Open weights
TextText

DeepSeek-V3

2024
DeepSeek

An open-weight MoE with 671B parameters, activating 37B per token

Model Open weights
TextText

GPT-4o

2024
OpenAI

A natively multimodal general model, with text, image and audio through one door

Model Closed
TextImageAudioTextAudio

Jamba

2024
AI21 Labs

An open-weight model that mixes a state-space model with a Transformer

Model Open weights
TextText

DBRX

2024
Databricks

Databricks’ open-weights MoE language model

Model Open weights
TextText

Mistral Large

2024
Mistral AI

The flagship commercial model from a European open-weights lab

Model Closed
TextText

Gemini

2024
Google DeepMind

One chat entry point that gathers search, office apps and a multimodal model

App Freemium
TextImageAudioTextImageAudio

Grok

2023
xAI

A chat model tied to a social platform’s live data

Model Closed
TextImageText

Yi

2023
01.AI

A bilingual Chinese–English open-weight model with very-long-context versions

Model Open weights
TextText

Hunyuan

2023
Tencent (Hunyuan)

Tencent’s general model family, with open-weight versions

Model Open weights
TextImageText

Doubao

2023
ByteDance (Seed)

ByteDance’s general chat model and application

Model Closed
TextImageText

Phi

2023
Microsoft

Small, efficient models built from curated data

Model Open weights
TextText

Step

2023
StepFun

A general model family aimed at multimodality and on-device use

Model Closed
TextImageText

Baichuan

2023
Baichuan AI

An open-weight general model aimed at Chinese-language use

Model Open weights
TextText

GLM

2023
Zhipu AI

A Chinese general model that began with autoregressive blank-infilling pretraining

Model Closed
TextText

Llama

2023
Meta AI (FAIR)

The model family that made the open-weight route mainstream

Model Open weights
TextText

ChatGPT

2022
OpenAI

The chat window that put a large language model in everyone’s hands

App Freemium
TextImageAudioTextImageAudio

Pangu

2020
Huawei

Huawei’s Pangu family of foundation models

Model Closed
TextText

Organizations involved

Typical uses

  • Solving maths and physics problems
  • Multi-step planning and scheduling
  • Code and data problems needing derivation
  • Reasoned choices under constraints

How it is evaluated

Math benchmark accuracy
Correct-answer rate on sets such as GSM8K, MATH, AIME
pass@k
Share of problems solved by at least one of k samples
Chain-of-thought faithfulness
Whether the written steps actually reflect how the answer was reached

Limits and hard parts

  • The written reasoning can be unfaithful: the answer comes first and a plausible rationale is added after
  • In long chains an early misstep propagates, yet the final answer still sounds assured
  • Deliberating at length on trivial questions inflates latency and cost

Concepts behind it