Imagen
Erzeugt fotorealistische Bilder mit kaskadierter Diffusion und großem Text-Encoder
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
Imagen is a text-to-image model Google announced in May 2022. Its key move is to use a frozen large language model (T5-XXL) as the text encoder and to generate with cascaded diffusion: a base image at low resolution is produced first, then a series of super-resolution diffusion models enlarge it step by step. The paper showed that scaling the text encoder improved image–text alignment more than scaling the image generator alone. Imagen did not release its weights and was not offered directly to the public.
Warum es wichtig ist
It showed that the size of the text encoder is a key lever for text–image alignment and pushed the cascaded-diffusion route to the frontier of photorealistic generation, shaping the design of later models.
Wichtige Eckdaten
- Text encoder
- Frozen T5-XXL
- Generation
- Cascaded diffusion (base + super-resolution)
- Output resolution
- 1024×1024
- Open weights
- No
Fähigkeiten
Verwandte Konzepte
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Vergleichbare Produkte
DALL·E 3
2023Schreibt eine lange Eingabe in eine detaillierte Beschreibung um und zeichnet dann das Bild
Stable Diffusion
2022Veröffentlichte Text-zu-Bild-Gewichte offen und klein genug für Consumer-GPUs
Midjourney
2022Ein Text-zu-Bild-Dienst, bekannt für seinen ästhetischen Stil
FLUX
2024Erzeugt hochauflösende Bilder mit einem Rectified-Flow-Transformer