DeepSeek-R1
A reasoning model trained with RL on chains of thought, its weights open under the MIT licence
WHAT IT IS
DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.
Why it matters
It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.
Key specs
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Open weights
- Yes (MIT licence)
- Type
- Reasoning model
Capabilities
Reasoning & Chain-of-Thought
Break a hard problem into intermediate steps
Text Generation
Continue a passage, one word at a time
Conversation & Instruction Following
Understand intent across turns and act on it
Code Generation
Write runnable code straight from a description
Related concepts
Reinforcement Learning from Human Feedback
When the good answer cannot be written as a formula, let humans stand in as the reward function
Prompting & Alignment
Making a model helpful, honest and harmless is harder than simply making it bigger
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Comparable products
o3
2025It reasons at length before answering, trading inference-time compute for steadier accuracy
DeepSeek-V3
2024An open-weight MoE with 671B parameters, activating 37B per token
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context