Text-to-Image
Turn a sentence into an image
WHAT THIS CAPABILITY MEANS
Maps a natural-language description directly to an image: text in, pixels out, with no annotation or reference image in between. Unlike image-to-image it has no input image at all, and unlike image editing it does not modify an existing picture but synthesises a new one from scratch.
How it is done
The dominant route is latent diffusion: the prompt is encoded into a conditioning vector, injected into a denoising network through cross-attention, denoised step by step in a compressed latent space, and decoded to pixels. Conditioning was later extended by classifier-free guidance, which amplifies prompt adherence by contrasting the conditioned and unconditioned directions. Training uses vast image–text pairs, and DALL·E 3 coupled prompt rewriting with generation in one pipeline, markedly improving adherence to long prompts.
Representative products
7DALL·E 3
2023Rewrites a long prompt into a detailed description, then draws the image
Midjourney
2022A text-to-image service known for its aesthetic style
Stable Diffusion
2022Released text-to-image weights openly and small enough to run on consumer GPUs
FLUX
2024Generates high-resolution images with a rectified-flow transformer
Imagen
2022Generates photorealistic images with cascaded diffusion and a large text encoder
Seedream
2024A text-to-image model with native high resolution and strong text rendering
Firefly
2023An image generation and editing tool aimed at creators
Organizations involved
Typical uses
- Concept design and storyboards
- Marketing assets and illustration
- Pre-visualisation for games and film
- Personalised avatars and wallpapers
How it is evaluated
- FID
- Distance to the real image distribution; lower is better
- CLIP score
- Semantic agreement between image and prompt
- Human-preference Elo
- Preference ranking from pairwise comparison
Limits and hard parts
- Counting, precise spatial relations and ordering remain unreliable
- Letters inside the image are frequently garbled, especially in long strings
- Hands, limbs and object-contact points break down, often needing several resamples
Concepts behind it
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Latent Diffusion & Conditional Control
Run diffusion not over pixels, but inside a compressed semantic space
Generative Models: An Overview
Discriminative models answer "what is this"; generative models answer "what should this look like"