Chuyển đến nội dung
Bản đồ AI

Imagen

Tạo ảnh chân thực bằng khuếch tán theo tầng và bộ mã hóa văn bản lớn

Google DeepMind Mô hình Đóng
đầu vàoVăn bảnHình ảnh

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

NÓ LÀ GÌ

Imagen is a text-to-image model Google announced in May 2022. Its key move is to use a frozen large language model (T5-XXL) as the text encoder and to generate with cascaded diffusion: a base image at low resolution is produced first, then a series of super-resolution diffusion models enlarge it step by step. The paper showed that scaling the text encoder improved image–text alignment more than scaling the image generator alone. Imagen did not release its weights and was not offered directly to the public.

Vì sao đáng ghi nhớ

It showed that the size of the text encoder is a key lever for text–image alignment and pushed the cascaded-diffusion route to the frontier of photorealistic generation, shaping the design of later models.

Thông số chính

Text encoder
Frozen T5-XXL
Generation
Cascaded diffusion (base + super-resolution)
Output resolution
1024×1024
Open weights
No

Năng lực liên quan

Khái niệm liên quan

Sản phẩm cùng loại