DeepSeek-R1
نموذج استدلال دُرِّب بالتعلم المعزز على سلاسل التفكير، وأوزانه متاحة تحت رخصة MIT
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
DeepSeek-R1 is a reasoning model DeepSeek released in January 2025 that produces a long chain of thought before answering. It trains the model with reinforcement learning to generate reasoning steps on its own, rather than relying on large sets of human-labelled rationales. R1 shares the architecture scale of the V3 family, releases its weights under the MIT licence and publishes a technical report on its training method. After release, many distilled versions of R1 appeared, transferring its reasoning ability into smaller models.
لماذا يستحق التذكّر
It fully open-sourced the weights of a frontier reasoning model under the MIT licence and published a reproducible RL training recipe; the wave of distilled versions that followed changed the cost structure of building one’s own reasoning capability.
المواصفات الأساسية
- Parameters
- 671B (MoE, 37B active)
- Context window
- 128K tokens
- Open weights
- Yes (MIT licence)
- Type
- Reasoning model
القدرات المرتبطة
المفاهيم ذات الصلة
التعلّم المعزّز من التغذية الراجعة البشرية
عندما يتعذّر كتابة «الإجابة الجيدة» كصيغة، دع البشر يقومون بدور دالة المكافأة
هندسة الأوامر والمواءمة
جعل النموذج مفيداً وصادقاً وغير ضارّ أصعب من مجرد تكبيره
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
منتجات منافسة
o3
2025يستنتج مطولًا قبل الإجابة، فيقايض حساب الاستدلال بدقة أكثر ثباتًا
DeepSeek-V3
2024نموذج MoE بأوزان مفتوحة بـ 671 مليار معامل، ينشّط 37 مليار لكل رمز
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
MiniMax-M
2025نموذج استدلال بأوزان مفتوحة بانتباه هجين وسياق مليون رمز