본문으로 건너뛰기
AI 도감

Imagen

계단식 확산과 대형 텍스트 인코더로 사실적인 이미지를 생성한다

Google DeepMind 모델 클로즈드 소스
입력텍스트이미지

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

무엇인가

Imagen is a text-to-image model Google announced in May 2022. Its key move is to use a frozen large language model (T5-XXL) as the text encoder and to generate with cascaded diffusion: a base image at low resolution is produced first, then a series of super-resolution diffusion models enlarge it step by step. The paper showed that scaling the text encoder improved image–text alignment more than scaling the image generator alone. Imagen did not release its weights and was not offered directly to the public.

기억할 만한 이유

It showed that the size of the text encoder is a key lever for text–image alignment and pushed the cascaded-diffusion route to the frontier of photorealistic generation, shaping the design of later models.

주요 사양

Text encoder
Frozen T5-XXL
Generation
Cascaded diffusion (base + super-resolution)
Output resolution
1024×1024
Open weights
No

소속 능력

관련 개념

동종 제품