يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما هو
Imagen is a text-to-image model Google announced in May 2022. Its key move is to use a frozen large language model (T5-XXL) as the text encoder and to generate with cascaded diffusion: a base image at low resolution is produced first, then a series of super-resolution diffusion models enlarge it step by step. The paper showed that scaling the text encoder improved image–text alignment more than scaling the image generator alone. Imagen did not release its weights and was not offered directly to the public.
لماذا يستحق التذكّر
It showed that the size of the text encoder is a key lever for text–image alignment and pushed the cascaded-diffusion route to the frontier of photorealistic generation, shaping the design of later models.
المواصفات الأساسية
- Text encoder
- Frozen T5-XXL
- Generation
- Cascaded diffusion (base + super-resolution)
- Output resolution
- 1024×1024
- Open weights
- No
القدرات المرتبطة
المفاهيم ذات الصلة
نماذج الانتشار
تعلّم ألف خطوة صغيرة لإزالة الضوضاء، فتستطيع بناء صورة من ضوضاء صافية
التوليد متعدد الوسائط
نموذج واحد يتعلّم الكلام والرسم والحركة، بل ونمذجة العالم ثلاثي الأبعاد
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
منتجات منافسة
DALL·E 3
2023يعيد صياغة الطلب الطويل إلى وصف مفصّل ثم يرسم الصورة
Stable Diffusion
2022أطلق أوزان النص إلى صورة مفتوحة وبحجم يعمل على كروت الرسوميات الاستهلاكية
Midjourney
2022خدمة نص إلى صورة تشتهر بذوقها الجمالي
FLUX
2024يولّد صورًا عالية الدقة بمحوّل تدفّق مُقوَّم