Generative Modelle im Überblick
Diskriminative Modelle beantworten "Was ist das?", generative "Wie sollte das aussehen?"
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
DEFINITION
A discriminative model learns the conditional distribution p(y|x) and maps an input directly to a label. A generative model learns the distribution of the data itself, p(x), so that it can sample new examples that follow the same distribution as the training data without duplicating it. The essential difference is what gets modelled: one draws a boundary, the other reconstructs the whole terrain. Generative models are also frequently asked to support conditional generation p(x|y), steering the output with text, a class label or a sketch.
Intuition
Picture an exam. A discriminative question only asks "is this photo a cat or a dog" — answer correctly and you are done. A generative question asks you to draw a cat convincing enough that a grader believes it is real. The first only needs to capture the difference between two classes; the second must know everything about cats — the shape of an ear, the direction of fur, the plausible range of poses. This is exactly why generation is often harder: what you must imitate is not a boundary but the entire shape of the data.
The divide between discrimination and generation: one answers "which class", the other must reconstruct "the whole distribution"
Training curves on one dataset: adversarial FID drops fast early but oscillates and rebounds later, while denoising diffusion declines smoothly and monotonically (illustrative magnitudes, lower is better)
- Adversarial (GAN)
- Denoising diffusion
Funktionsweise
- 01
Fix the distribution to model
First decide the target: reproduce the whole dataset unconditionally, p(x), or generate under a condition such as text or a class label, p(x|y). This choice fixes the model’s output interface and where its training signal comes from.
- 02
Choose how to represent the density
Explicit-density models (autoregressive, VAE, diffusion) write down or approximate a functional form for p(x) and can be trained by likelihood; implicit-density models (GAN) never give a probability, only a sampler, and learn from an adversarial real-or-fake signal. The former has a clear objective but is constrained; the latter is flexible but hard to train.
- 03
Optimise with likelihood or an adversarial signal
Autoregressive models factorise the joint probability into a product of per-step predictions; VAEs optimise a lower bound on the data likelihood; diffusion splits generation into hundreds of denoising steps; GANs let a discriminator serve as a learned loss. Four routes, four mathematical formulations of "looking real".
- 04
Measure quality with metrics and human judgement
FID compares the generated and real distributions in Inception feature space, and IS scores the sharpness and diversity of individual samples — both are only proxies. The trustworthy judge remains human preference, whether by survey or pairwise arena voting.
Four principal routes: each gives a different mathematical definition of "looking real"
Representative FID magnitudes on ImageNet 256×256 by route (lower is better; autoregressive as VQGAN+Transformer, autoencoder as VQGAN, adversarial as BigGAN-deep, diffusion as DiT-XL/2)
Anwendungsfelder
- Image and video generation: synthesising new visual content from text, sketches or reference images
- Speech and music synthesis: text-to-speech, singing-voice synthesis, music continuation
- Data augmentation and simulation: synthesising samples for rare classes, or synthetic scenes for autonomous driving
- Scientific design: generating candidate molecules, protein structures or chip layouts
Häufige Missverständnisse
- Realistic output does not imply the right distribution was learned. A model that is sharp but low in diversity — mode collapse — can look good on FID while covering only a small part of the data.
- FID and IS depend on a fixed feature network; swap the extractor and the ranking can flip. They correlate with human intuition, they do not equal it.
- There is no answer key for generative output. One prompt admits countless reasonable results, so evaluation must separate "implausible" from merely "unusual".
Schlüsselbegriffe
- Explicit density
- A model that writes down or approximates p(x), e.g. autoregressive or diffusion
- Implicit density
- A model that offers only a sampler, not a probability, e.g. a GAN
- Mode collapse
- When a generator covers only a few modes of the data distribution
- FID
- Fréchet distance between generated and real distributions in Inception feature space; lower is better