이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
무엇인가
DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.
기억할 만한 이유
It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.
주요 사양
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Open weights
- Yes (MIT licence)
- Type
- Reasoning model