확산 모델
작은 잡음 제거를 천 번 배우면 순수 잡음에서 이미지를 만들어낸다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
A diffusion model defines a forward process that gradually adds noise until the data becomes approximately Gaussian noise, then trains a network to invert that process, recovering data from noise step by step. Generation starts from pure noise and repeatedly applies the learned denoising steps. DDPM casts this as a discrete Markov chain requiring around a thousand sampling steps; deterministic samplers such as DDIM reach comparable quality in a few dozen.
직관적 이해
Instead of demanding that a pen draw a picture in one stroke, train it to clean a photograph blurred by noise. Removing a little noise is an easy task to learn; chaining a thousand such clean-ups amounts to generating an image out of pure noise. Breaking a hard problem into a chain of small ones is precisely why diffusion is easier to train than a GAN.
Forward noising turns an image into noise step by step (top row); reverse denoising recovers the image from noise (bottom row)
FID versus step count for deterministic samplers: DDIM needs nearly a hundred steps to saturate, while higher-order solvers such as DPM-Solver reach comparable quality in twenty to thirty (CIFAR-10 magnitudes)
작동 원리
- 01
Forward noising
According to a fixed noise schedule, Gaussian noise is added to the data step by step until it is almost pure noise. This process has no learnable parameters; it is purely destructive.
- 02
Learn the reverse denoising
Train a network to predict the noise added at each step, or equivalently to predict the clean image. Because the forward process is known, the target is deterministic given the noisy image and the timestep — reducing training to a stable regression problem.
- 03
Reverse sampling
Starting from standard Gaussian noise, the network estimates and removes noise at each step, yielding progressively cleaner samples. Every DDPM step injects randomness, whereas DDIM takes a deterministic path, allowing larger strides and far fewer steps.
- 04
Classifier-free guidance
During training the condition (e.g. text) is randomly dropped, forcing one network to handle both conditional and unconditional generation. At sampling time the two predictions are extrapolated by a guidance scale, giving a continuous knob from diversity to prompt fidelity: a higher scale hews closer to the prompt but reduces variety.
핵심 수식
xₜ = √(ᾱₜ) · x₀ + √(1 − ᾱₜ) · ε, ε ~ N(0, I)응용 분야
- Text-to-image and image editing: repainting, local modification and stylisation
- High-quality image and video synthesis, including super-resolution
- Speech and audio generation, and molecular or protein structure generation
- As a simulator: reconstructing signals from noise for scientific and engineering inverse problems
흔한 오해
- Diffusion is slow only with early samplers. DDIM, DPM-Solver, distillation and consistency models have compressed sampling to a dozen steps or even one.
- A larger guidance scale does not simply mean closer to the prompt. Too high causes oversaturation and rigid structure while reducing diversity; it is not strictly better.
- Running directly in pixel space is expensive, especially at high resolution — which is exactly the motivation for latent diffusion.
핵심 용어
- Noise schedule
- The timetable of noise added per step, described by βₜ or ᾱₜ
- DDPM
- Discrete Markov diffusion, typically needing a thousand sampling steps
- DDIM
- Deterministic sampling achieving comparable quality in a few dozen steps
- Classifier-free guidance
- Extrapolating between conditional and unconditional predictions to control prompt fidelity