텍스트-이미지 생성
한 문장 설명으로 이미지 한 장을 만든다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.
기술적으로 구현하는 방법
The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.
대표 제품
7DALL·E 3
2023긴 프롬프트를 상세한 설명으로 다시 쓴 뒤 이미지를 생성한다
Midjourney
2022미적 스타일로 알려진 텍스트-이미지 서비스
Stable Diffusion
2022텍스트-이미지 가중치를 공개하고 소비자용 GPU에서 돌아가게 만들었다
FLUX
2024정류 흐름 트랜스포머로 고해상도 이미지를 생성한다
Imagen
2022계단식 확산과 대형 텍스트 인코더로 사실적인 이미지를 생성한다
Seedream
2024네이티브 고해상도와 뛰어난 문자 렌더링의 텍스트-이미지 모델
Firefly
2023크리에이터를 위한 이미지 생성·편집 도구
관련 기관
대표적 용도
- Concept design and storyboards
- Marketing assets and illustration
- Pre-visualisation for games and film
- Personalised avatars and wallpapers
성능을 평가하는 방법
- FID
- Distance to the real image distribution; lower is better
- CLIP score
- Semantic agreement between image and prompt
- Human-preference Elo
- Preference ranking from pairwise comparison
경계와 난점
- Counting, precise spatial relations and ordering remain unreliable
- Letters inside the image are frequently garbled, especially in long strings
- Hands, limbs and object-contact points break down, often needing several resamples