DeepSeek-R1
तर्क-श्रृंखला पर RL से प्रशिक्षित तर्क मॉडल, भार MIT लाइसेंस के तहत खुले
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.
यह क्यों महत्वपूर्ण है
It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.
मुख्य विशिष्टताएँ
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Open weights
- Yes (MIT licence)
- Type
- Reasoning model
संबंधित क्षमताएँ
तर्क और विचार-शृंखला
कठिन समस्या को मध्यवर्ती चरणों में बाँटकर हल करना
टेक्स्ट जनरेशन
पहले के पाठ को शब्द-दर-शब्द आगे बढ़ाता है
संवाद और निर्देश-पालन
बातचीत में मंशा समझकर उस पर अमल करना
कोड जनरेशन
विवरण से चलने योग्य कोड लिखना
संबंधित अवधारणाएँ
मानव प्रतिक्रिया से सुदृढ़ीकरण अधिगम
जब «अच्छा उत्तर» सूत्र में न लिखा जा सके, मनुष्य को पुरस्कार फलन बनने दें
प्रॉम्प्ट इंजीनियरिंग और अलाइनमेंट
मॉडल को उपयोगी, ईमानदार और हानिरहित बनाना उसे केवल बड़ा करने से कठिन है
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
समान उत्पाद
o3
2025उत्तर देने से पहले लंबी तर्क-श्रृंखला, अनुमान के समय गणना से अधिक सटीकता
DeepSeek-V3
2024671B पैरामीटर वाला ओपन-वेट MoE, प्रति टोकन केवल 37B सक्रिय
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
MiniMax-M
2025हाइब्रिड अटेंशन और दस लाख टोकन संदर्भ वाला ओपन-वेट तर्क मॉडल