Imagen
Generates photorealistic images with cascaded diffusion and a large text encoder
WHAT IT IS
Imagen is a text-to-image model Google announced in May 2022. Its key move is to use a frozen large language model (T5-XXL) as the text encoder and to generate with cascaded diffusion: a base image at low resolution is produced first, then a series of super-resolution diffusion models enlarge it step by step. The paper showed that scaling the text encoder improved image–text alignment more than scaling the image generator alone. Imagen did not release its weights and was not offered directly to the public.
Why it matters
It showed that the size of the text encoder is a key lever for text–image alignment and pushed the cascaded-diffusion route to the frontier of photorealistic generation, shaping the design of later models.
Key specs
- Text encoder
- Frozen T5-XXL
- Generation
- Cascaded diffusion (base + super-resolution)
- Output resolution
- 1024×1024
- Open weights
- No
Capabilities
Related concepts
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Comparable products
DALL·E 3
2023Rewrites a long prompt into a detailed description, then draws the image
Stable Diffusion
2022Released text-to-image weights openly and small enough to run on consumer GPUs
Midjourney
2022A text-to-image service known for its aesthetic style
FLUX
2024Generates high-resolution images with a rectified-flow transformer