Autoencodeurs et VAE
Comprimer l’information dans un goulot, puis la laisser repousser
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
DÉFINITION
An autoencoder is a pair of networks: an encoder compresses the input into a low-dimensional representation, and a decoder tries to reconstruct the input from it, with reconstruction error as the training objective. A variational autoencoder (VAE) replaces that representation — a deterministic point — with a probability distribution, usually an axis-aligned Gaussian, and adds a penalty that pulls the whole latent space toward a standard normal distribution, so that new samples can be drawn by random sampling.
Intuition
An autoencoder trained only on compress-and-restore is like memorising an entire phone book: it knows every number but has extracted no pattern, so it cannot invent a new one. A VAE demands more: not only must reconstruction be accurate, the latent space itself must be orderly — nearby codes decode to nearby things, and every region must be habitable. Then an arbitrary point still decodes to something plausible. The cost of this constraint is a slightly blurrier reconstruction than a plain autoencoder.
The encode–decode loop of an autoencoder; a VAE replaces the deterministic vector at the bottleneck with a samplable Gaussian
The KL term pulls each latent dimension toward a standard normal: once training converges, the latent marginal approaches N(0, 1), so random points still decode to plausible samples
Fonctionnement
- 01
Encode into the bottleneck
The encoder reduces dimensionality layer by layer into a latent vector far smaller than the input. The bottleneck matters because it forces the model to discard detail and keep reusable structure — the origin of both dimensionality reduction and feature extraction.
- 02
Decode back to the input
The decoder expands the latent vector back to the input’s size, and the loss is typically per-pixel mean squared error or binary cross-entropy. Reconstruction loss cares only about matching the original, not about whether the latent space is well organised.
- 03
VAE’s probabilistic twist
The encoder no longer outputs a point but the mean and variance of the latent variable, defining a Gaussian; decoding samples from it before feeding the decoder. The reparameterisation trick writes sampling as mean plus standard deviation times noise, letting gradients flow back through a random operation.
- 04
ELBO: a tug-of-war
The training objective, the evidence lower bound (ELBO), has two terms: a reconstruction term that wants accurate decoding, and a regulariser (KL divergence) that pulls the encoded distributions toward a standard normal. The first pushes codes apart to stay distinctive, the second pulls them together to stay orderly, and their balance sets the sharpness and diversity of generated samples.
Formule clé
L = E_q[log p(x|z)] − KL( q(z|x) ‖ p(z) )Où c'est utilisé
- Representation learning and dimensionality reduction: using the bottleneck for feature extraction, visualisation and denoising
- Generative modelling: sampling the latent space to synthesise images, molecules and sequences
- Anomaly detection: anomalies reconstruct with markedly higher error than normal samples
- As a front end for other generative models: the encoder and decoder inside latent diffusion are exactly an autoencoder
Idées fausses courantes
- VAE samples are often blurry. A per-pixel reconstruction loss averages over several plausible answers, yielding an average image that resembles none of them; diffusion and GANs fill precisely this gap.
- A plain autoencoder is not a generative model. Its latent space has no samplable structure, so an arbitrary point usually decodes to something meaningless.
- A smaller bottleneck is not automatically better. Too narrow loses necessary information and reconstruction collapses; too wide degenerates into an identity map that learns nothing useful.
Termes clés
- Bottleneck
- The low-dimensional layer holding the latent code, limiting its bandwidth
- Reparameterisation
- Writing sampling as a deterministic transform plus external noise so gradients flow
- KL divergence
- Measures how far the encoded distribution deviates from a standard normal; acts as a regulariser
- ELBO
- A lower bound on the log-likelihood: the reconstruction term minus the KL term; a VAE’s actual objective