본문으로 건너뛰기
AI 도감

오토인코더와 변분 오토인코더

정보를 병목으로 압축했다가 다시 길러낸다

07 생성형 AI중급이 영역의 2번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

An autoencoder is a pair of networks: an encoder compresses the input into a low-dimensional representation, and a decoder tries to reconstruct the input from it, with reconstruction error as the training objective. A variational autoencoder (VAE) replaces that representation — a deterministic point — with a probability distribution, usually an axis-aligned Gaussian, and adds a penalty that pulls the whole latent space toward a standard normal distribution, so that new samples can be drawn by random sampling.

직관적 이해

An autoencoder trained only on compress-and-restore is like memorising an entire phone book: it knows every number but has extracted no pattern, so it cannot invent a new one. A VAE demands more: not only must reconstruction be accurate, the latent space itself must be orderly — nearby codes decode to nearby things, and every region must be habitable. Then an arbitrary point still decodes to something plausible. The cost of this constraint is a slightly blurrier reconstruction than a plain autoencoder.

그림 1

The encode–decode loop of an autoencoder; a VAE replaces the deterministic vector at the bottleneck with a samplable Gaussian

그림 2

The KL term pulls each latent dimension toward a standard normal: once training converges, the latent marginal approaches N(0, 1), so random points still decode to plausible samples

작동 원리

  1. 01

    Encode into the bottleneck

    The encoder reduces dimensionality layer by layer into a latent vector far smaller than the input. The bottleneck matters because it forces the model to discard detail and keep reusable structure — the origin of both dimensionality reduction and feature extraction.

  2. 02

    Decode back to the input

    The decoder expands the latent vector back to the input’s size, and the loss is typically per-pixel mean squared error or binary cross-entropy. Reconstruction loss cares only about matching the original, not about whether the latent space is well organised.

  3. 03

    VAE’s probabilistic twist

    The encoder no longer outputs a point but the mean and variance of the latent variable, defining a Gaussian; decoding samples from it before feeding the decoder. The reparameterisation trick writes sampling as mean plus standard deviation times noise, letting gradients flow back through a random operation.

  4. 04

    ELBO: a tug-of-war

    The training objective, the evidence lower bound (ELBO), has two terms: a reconstruction term that wants accurate decoding, and a regulariser (KL divergence) that pulls the encoded distributions toward a standard normal. The first pushes codes apart to stay distinctive, the second pulls them together to stay orderly, and their balance sets the sharpness and diversity of generated samples.

핵심 수식

L = E_q[log p(x|z)] − KL( q(z|x) ‖ p(z) )
The evidence lower bound (ELBO): the first term is the reconstruction likelihood (decode accurately), the second is the KL regulariser (stay close to the prior p(z)); a VAE maximises this bound.

응용 분야

  • Representation learning and dimensionality reduction: using the bottleneck for feature extraction, visualisation and denoising
  • Generative modelling: sampling the latent space to synthesise images, molecules and sequences
  • Anomaly detection: anomalies reconstruct with markedly higher error than normal samples
  • As a front end for other generative models: the encoder and decoder inside latent diffusion are exactly an autoencoder

흔한 오해

  • VAE samples are often blurry. A per-pixel reconstruction loss averages over several plausible answers, yielding an average image that resembles none of them; diffusion and GANs fill precisely this gap.
  • A plain autoencoder is not a generative model. Its latent space has no samplable structure, so an arbitrary point usually decodes to something meaningless.
  • A smaller bottleneck is not automatically better. Too narrow loses necessary information and reconstruction collapses; too wide degenerates into an identity map that learns nothing useful.

핵심 용어

Bottleneck
The low-dimensional layer holding the latent code, limiting its bandwidth
Reparameterisation
Writing sampling as a deterministic transform plus external noise so gradients flow
KL divergence
Measures how far the encoded distribution deviates from a standard normal; acts as a regulariser
ELBO
A lower bound on the log-likelihood: the reconstruction term minus the KL term; a VAE’s actual objective

참고문헌