본문으로 건너뛰기
AI 도감
07 생성형 AI중급이 영역의 4번째 항목

확산 모델

작은 잡음 제거를 천 번 배우면 순수 잡음에서 이미지를 만들어낸다

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

A diffusion model defines a forward process that gradually adds noise until the data becomes approximately Gaussian noise, then trains a network to invert that process, recovering data from noise step by step. Generation starts from pure noise and repeatedly applies the learned denoising steps. DDPM casts this as a discrete Markov chain requiring around a thousand sampling steps; deterministic samplers such as DDIM reach comparable quality in a few dozen.

직관적 이해

Instead of demanding that a pen draw a picture in one stroke, train it to clean a photograph blurred by noise. Removing a little noise is an easy task to learn; chaining a thousand such clean-ups amounts to generating an image out of pure noise. Breaking a hard problem into a chain of small ones is precisely why diffusion is easier to train than a GAN.

그림 1

Forward noising turns an image into noise step by step (top row); reverse denoising recovers the image from noise (bottom row)

그림 2

FID versus step count for deterministic samplers: DDIM needs nearly a hundred steps to saturate, while higher-order solvers such as DPM-Solver reach comparable quality in twenty to thirty (CIFAR-10 magnitudes)

작동 원리

  1. 01

    Forward noising

    According to a fixed noise schedule, Gaussian noise is added to the data step by step until it is almost pure noise. This process has no learnable parameters; it is purely destructive.

  2. 02

    Learn the reverse denoising

    Train a network to predict the noise added at each step, or equivalently to predict the clean image. Because the forward process is known, the target is deterministic given the noisy image and the timestep — reducing training to a stable regression problem.

  3. 03

    Reverse sampling

    Starting from standard Gaussian noise, the network estimates and removes noise at each step, yielding progressively cleaner samples. Every DDPM step injects randomness, whereas DDIM takes a deterministic path, allowing larger strides and far fewer steps.

  4. 04

    Classifier-free guidance

    During training the condition (e.g. text) is randomly dropped, forcing one network to handle both conditional and unconditional generation. At sampling time the two predictions are extrapolated by a guidance scale, giving a continuous knob from diversity to prompt fidelity: a higher scale hews closer to the prompt but reduces variety.

핵심 수식

xₜ = √(ᾱₜ) · x₀ + √(1 − ᾱₜ) · ε, ε ~ N(0, I)
Closed-form noising: the noisy image at any timestep t follows directly from the clean image and one standard Gaussian noise, which is why the training target is deterministic.

응용 분야

  • Text-to-image and image editing: repainting, local modification and stylisation
  • High-quality image and video synthesis, including super-resolution
  • Speech and audio generation, and molecular or protein structure generation
  • As a simulator: reconstructing signals from noise for scientific and engineering inverse problems

흔한 오해

  • Diffusion is slow only with early samplers. DDIM, DPM-Solver, distillation and consistency models have compressed sampling to a dozen steps or even one.
  • A larger guidance scale does not simply mean closer to the prompt. Too high causes oversaturation and rigid structure while reducing diversity; it is not strictly better.
  • Running directly in pixel space is expensive, especially at high resolution — which is exactly the motivation for latent diffusion.

핵심 용어

Noise schedule
The timetable of noise added per step, described by βₜ or ᾱₜ
DDPM
Discrete Markov diffusion, typically needing a thousand sampling steps
DDIM
Deterministic sampling achieving comparable quality in a few dozen steps
Classifier-free guidance
Extrapolating between conditional and unconditional predictions to control prompt fidelity

참고문헌