DeepSeek-V3
نموذج MoE بأوزان مفتوحة بـ 671 مليار معامل، ينشّط 37 مليار لكل رمز
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
DeepSeek-V3 is an open-weight model DeepSeek released in December 2024. It uses a mixture-of-experts architecture with 671B total parameters, activating about 37B per token. DeepSeek was founded in Hangzhou in 2023 by Liang Wenfeng and is known for publishing technical details and data openly. V3 approaches the closed frontier of its time on several benchmarks, while the training compute and cost disclosed in its technical report are far below the usual level for models of this size. It supports a context window in the 128K range, with weights downloadable under a permissive licence.
لماذا يستحق التذكّر
A 671B-parameter MoE that activates only 37B per token, trained at the low cost reported in its technical report — evidence that open-weight models can approach the closed frontier of the same period.
المواصفات الأساسية
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Training cost
- About US$5.58M (per the official technical report)
- Open weights
- Yes
القدرات المرتبطة
المفاهيم ذات الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
التدريب المسبق والضبط الدقيق
تعلّم اللغة أولاً من نصوص ضخمة بلا وسوم ثم التخصص ببيانات قليلة — أكثر النماذج كفاءةً في البيانات
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
منتجات منافسة
DeepSeek-R1
2025نموذج استدلال دُرِّب بالتعلم المعزز على سلاسل التفكير، وأوزانه متاحة تحت رخصة MIT
Llama
2023عائلة النماذج التي جعلت مسار الأوزان المفتوحة سائدًا
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
Mistral Large
2024النموذج التجاري الرائد من مختبر أوروبي للأوزان المفتوحة
Baichuan
2023نموذج عام بأوزان مفتوحة موجّه للاستخدام بالصينية