本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.
なぜ覚えておく価値があるか
It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.
主な仕様
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Open weights
- Yes (MIT licence)
- Type
- Reasoning model
対応する能力
関連する概念
同種の製品
o3
2025答える前に長い推論を重ね、推論時の計算で正答率を上げる
DeepSeek-V3
2024総パラメータ 671B、1 トークンあたり 37B だけを活性化するオープンウェイト MoE
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
MiniMax-M
2025ハイブリッド注意と100万トークン文脈をもつオープンウェイトの推論モデル