यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्या है
DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.
यह क्यों महत्वपूर्ण है
A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.
मुख्य विशिष्टताएँ
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Training cost
- About US$5.58M (per the official technical report)
- Open weights
- Yes
संबंधित क्षमताएँ
टेक्स्ट जनरेशन
पहले के पाठ को शब्द-दर-शब्द आगे बढ़ाता है
संवाद और निर्देश-पालन
बातचीत में मंशा समझकर उस पर अमल करना
तर्क और विचार-शृंखला
कठिन समस्या को मध्यवर्ती चरणों में बाँटकर हल करना
कोड जनरेशन
विवरण से चलने योग्य कोड लिखना
संबंधित अवधारणाएँ
Transformer आर्किटेक्चर
शब्द-दर-शब्द relay की जगह वह कक्ष जहाँ सब एक साथ बोलते हैं, जिससे दूर की निर्भरताएँ एक कदम पर आ जाती हैं
प्री-ट्रेनिंग और फाइन-ट्यूनिंग
पहले विशाल अलेबल पाठ से भाषा सीखना, फिर थोड़े डेटा से विशेषज्ञ बनना — आधुनिक एआई का सबसे डेटा-कुशल प्रतिमान
अटेंशन तंत्र
हर स्थान बाकी सभी स्थानों को सीधे देख सकता है और प्रासंगिकता के अनुसार ध्यान बाँट सकता है
समान उत्पाद
DeepSeek-R1
2025तर्क-श्रृंखला पर RL से प्रशिक्षित तर्क मॉडल, भार MIT लाइसेंस के तहत खुले
Llama
2023खुले भार के रास्ते को मुख्यधारा बनाने वाला मॉडल परिवार
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Mistral Large
2024यूरोपीय ओपन-वेट लैब का फ़्लैगशिप वाणिज्यिक मॉडल
Baichuan
2023चीनी-भाषा उपयोग के लिए ओपन-वेट सामान्य मॉडल