본문으로 건너뛰기
AI 도감

텍스트-이미지 생성

한 문장 설명으로 이미지 한 장을 만든다

이미지 생성과 편집입문 #18
입력텍스트이미지

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.

기술적으로 구현하는 방법

The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.

대표 제품

7

관련 기관

대표적 용도

  • Concept design and storyboards
  • Marketing assets and illustration
  • Pre-visualisation for games and film
  • Personalised avatars and wallpapers

성능을 평가하는 방법

FID
Distance to the real image distribution; lower is better
CLIP score
Semantic agreement between image and prompt
Human-preference Elo
Preference ranking from pairwise comparison

경계와 난점

  • Counting, precise spatial relations and ordering remain unreliable
  • Letters inside the image are frequently garbled, especially in long strings
  • Hands, limbs and object-contact points break down, often needing several resamples

뒤에 있는 개념