Chuyển đến nội dung
Bản đồ AI

Khuếch tán trong không gian ẩn và điều khiển có điều kiện

Khuếch tán không trên điểm ảnh, mà trong một không gian ngữ nghĩa đã nén

07 AI tạo sinhChuyên giaMục từ thứ 5 trong lĩnh vực

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

Latent diffusion first compresses an image into a low-dimensional latent representation with an autoencoder, runs the diffusion process inside that latent space, and finally decodes back to pixels. Stable Diffusion exemplifies the paradigm: a variational autoencoder converts between pixels and latents, a U-Net denoises in latent space, and a text encoder injects the prompt through cross-attention. ControlNet, LoRA and similar methods then attach to or lightly fine-tune the frozen backbone to add structural or stylistic control.

Trực giác

Diffusing over raw pixels is expensive: a 512×512 image is 262,000 numbers, iterated dozens of times. Yet the gist of an image is far smaller — colours, composition and object relations all fit in a more compact code. SD first summarises the image into an outline, creates within that outline, and finally expands it back into prose. The saved compute is orders of magnitude, with essentially no loss in quality.

Hình 1

The three-part structure of Stable Diffusion: pixels and text are encoded separately, all resampling happens in latent space, and a decoder restores pixels

That all resampling happens in latent space is the source of the compute saving and of high-resolution scalability.
Hình 2

Human-preference Elo of leading text-to-image models (illustrative magnitudes from public text-to-image arenas; values shift as leaderboards update)

Cách hoạt động

  1. 01

    Compress with a perceptual autoencoder

    An autoencoder trained with perceptual and adversarial losses compresses an image to roughly one-eighth per side (1/64 the area) in latent space. The downsampling factor is a trade-off: more compression saves compute, but excessive compression discards detail.

  2. 02

    Denoise in latent space

    A U-Net performs diffusion on the latent representation, taking the noisy latent and the timestep and predicting the noise. All resampling happens in low dimensions, cutting both training and inference cost substantially.

  3. 03

    Inject text via cross-attention

    A text encoder turns the prompt into a sequence of vectors; at every U-Net layer, cross-attention lets image features query those text vectors, so the prompt participates at each denoising stage rather than as a single upfront condition.

  4. 04

    Conditional control on a frozen backbone

    ControlNet duplicates the encoder side and adds a zero-convolution bypass, guiding structure with edges, depth or pose maps; LoRA trains only low-rank incremental matrices, adapting new styles with a small footprint. Neither alters the backbone’s original weights.

Hình 3

Cross-attention: every image-feature position queries all text tokens, receiving the semantic condition at each denoising step

Hình 4

Latent diffusion moves resampling onto a representation roughly 1/64 the size, cutting training and sampling compute to a fraction of pixel-space diffusion (illustrative magnitudes)

  • Pixel-space diffusion
  • Latent diffusion

Ứng dụng

  • Text-to-image, image editing and inpainting or outpainting
  • Structure-controlled generation: pose, depth, line art or segmentation to image
  • Personalisation: customising subjects or styles with LoRA and DreamBooth
  • Efficient video and 3D generation: the same paradigm extended to time and space

Hiểu lầm thường gặp

  • Latent space is not a free lunch. If the autoencoder loses information, no amount of diffusion can restore it, and high-resolution detail is especially fragile.
  • LoRA adapts by addition, not replacement. Style or subject lives in the low-rank increment and can interfere with new prompts; too large a weight also corrupts existing abilities.
  • ControlNet strength and prompt must work together. Too strong a control skips semantics and merely copies the condition map; too weak makes it inert.

Thuật ngữ chính

Latent space
The low-dimensional representation space produced by the autoencoder
Cross-attention
The attention mechanism letting image features query text vectors
ControlNet
A bypass network guiding structure from a condition map
LoRA
Low-rank adaptation increments for low-cost customisation

Đọc thêm