يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
Gemini is Google DeepMind’s multimodal model family, announced in December 2023 and designed from the start to handle text, image, audio and video together. It is integrated across Google’s existing products, serving Search, Workspace and cloud platforms at once. The Gemini 1.5 generation extended the context window to the million-token range, enough to read long documents, code bases or hours of video in one pass. It is both a model and an application brand, spanning lightweight on-device to large cloud-scale versions.
لماذا يستحق التذكّر
By pushing the context window to the million-token range and demonstrating it publicly, it reset expectations for how much a model can read in one pass and made multimodal input standard for general models.
المواصفات الأساسية
- Context window
- 1M tokens (Gemini 1.5 Pro)
- Modality
- Text, image, audio, video in; text out
- Released
- 2023-12
- Open weights
- No
القدرات المرتبطة
المفاهيم ذات الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة
التوليد متعدد الوسائط
نموذج واحد يتعلّم الكلام والرسم والحركة، بل ونمذجة العالم ثلاثي الأبعاد
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
منتجات منافسة
GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Claude
2023نموذج محادثة عام يشتهر بسياقه الطويل ومواءمته الأمنية
Llama
2023عائلة النماذج التي جعلت مسار الأوزان المفتوحة سائدًا
o3
2025يستنتج مطولًا قبل الإجابة، فيقايض حساب الاستدلال بدقة أكثر ثباتًا
Grok
2023نموذج محادثة مرتبط ببيانات منصة اجتماعية