يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
Jamba is an open-weight model AI21 Labs released in March 2024, interleaving Mamba state-space layers with Transformer attention layers inside one mixture-of-experts architecture. AI21 Labs was founded in Tel Aviv in 2017 by Amnon Shashua and others. This hybrid design holds long context while cutting memory and compute relative to a pure-attention model. Jamba ships as a 52B-parameter MoE that activates about 12B per token, with a context window in the 256K range.
لماذا يستحق التذكّر
It replaced most attention layers with state-space layers, testing whether a non-pure-Transformer architecture can save compute at long context — a concrete instance of hybrid architectures entering open-weight models.
المواصفات الأساسية
- Parameters
- 52B (MoE, ~12B active)
- Context window
- 256K tokens
- Architecture
- Mamba state-space layers interleaved with Transformer layers
- Open weights
- Yes
القدرات المرتبطة
المفاهيم ذات الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
التدريب المسبق والضبط الدقيق
تعلّم اللغة أولاً من نصوص ضخمة بلا وسوم ثم التخصص ببيانات قليلة — أكثر النماذج كفاءةً في البيانات