Diffusion latente et contrôle conditionnel
Diffuser non pas sur les pixels, mais dans un espace sémantique comprimé
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
DÉFINITION
Latent diffusion first compresses an image into a low-dimensional latent representation with an autoencoder, runs the diffusion process inside that latent space, and finally decodes back to pixels. Stable Diffusion exemplifies the paradigm: a variational autoencoder converts between pixels and latents, a U-Net denoises in latent space, and a text encoder injects the prompt through cross-attention. ControlNet, LoRA and similar methods then attach to or lightly fine-tune the frozen backbone to add structural or stylistic control.
Intuition
Diffusing over raw pixels is expensive: a 512×512 image is 262,000 numbers, iterated dozens of times. Yet the gist of an image is far smaller — colours, composition and object relations all fit in a more compact code. SD first summarises the image into an outline, creates within that outline, and finally expands it back into prose. The saved compute is orders of magnitude, with essentially no loss in quality.
The three-part structure of Stable Diffusion: pixels and text are encoded separately, all resampling happens in latent space, and a decoder restores pixels
Human-preference Elo of leading text-to-image models (illustrative magnitudes from public text-to-image arenas; values shift as leaderboards update)
Fonctionnement
- 01
Compress with a perceptual autoencoder
An autoencoder trained with perceptual and adversarial losses compresses an image to roughly one-eighth per side (1/64 the area) in latent space. The downsampling factor is a trade-off: more compression saves compute, but excessive compression discards detail.
- 02
Denoise in latent space
A U-Net performs diffusion on the latent representation, taking the noisy latent and the timestep and predicting the noise. All resampling happens in low dimensions, cutting both training and inference cost substantially.
- 03
Inject text via cross-attention
A text encoder turns the prompt into a sequence of vectors; at every U-Net layer, cross-attention lets image features query those text vectors, so the prompt participates at each denoising stage rather than as a single upfront condition.
- 04
Conditional control on a frozen backbone
ControlNet duplicates the encoder side and adds a zero-convolution bypass, guiding structure with edges, depth or pose maps; LoRA trains only low-rank incremental matrices, adapting new styles with a small footprint. Neither alters the backbone’s original weights.
Cross-attention: every image-feature position queries all text tokens, receiving the semantic condition at each denoising step
Latent diffusion moves resampling onto a representation roughly 1/64 the size, cutting training and sampling compute to a fraction of pixel-space diffusion (illustrative magnitudes)
- Pixel-space diffusion
- Latent diffusion
Où c'est utilisé
- Text-to-image, image editing and inpainting or outpainting
- Structure-controlled generation: pose, depth, line art or segmentation to image
- Personalisation: customising subjects or styles with LoRA and DreamBooth
- Efficient video and 3D generation: the same paradigm extended to time and space
Idées fausses courantes
- Latent space is not a free lunch. If the autoencoder loses information, no amount of diffusion can restore it, and high-resolution detail is especially fragile.
- LoRA adapts by addition, not replacement. Style or subject lives in the low-rank increment and can interfere with new prompts; too large a weight also corrupts existing abilities.
- ControlNet strength and prompt must work together. Too strong a control skips semantics and merely copies the condition map; too weak makes it inert.
Termes clés
- Latent space
- The low-dimensional representation space produced by the autoencoder
- Cross-attention
- The attention mechanism letting image features query text vectors
- ControlNet
- A bypass network guiding structure from a condition map
- LoRA
- Low-rank adaptation increments for low-cost customisation