本文へスキップ
AI図鑑

生成モデルの全体像

判別モデルは「これは何か」に、生成モデルは「どう見えるべきか」に答える

07 生成 AI初級この領域の第 1 項目

本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。

定義

A discriminative model learns the conditional distribution p(y|x) and maps an input directly to a label. A generative model learns the distribution of the data itself, p(x), so that it can sample new examples that follow the same distribution as the training data without duplicating it. The essential difference is what gets modelled: one draws a boundary, the other reconstructs the whole terrain. Generative models are also frequently asked to support conditional generation p(x|y), steering the output with text, a class label or a sketch.

直観的な理解

Picture an exam. A discriminative question only asks "is this photo a cat or a dog" — answer correctly and you are done. A generative question asks you to draw a cat convincing enough that a grader believes it is real. The first only needs to capture the difference between two classes; the second must know everything about cats — the shape of an ear, the direction of fur, the plausible range of poses. This is exactly why generation is often harder: what you must imitate is not a boundary but the entire shape of the data.

図 1

The divide between discrimination and generation: one answers "which class", the other must reconstruct "the whole distribution"

図 2

Training curves on one dataset: adversarial FID drops fast early but oscillates and rebounds later, while denoising diffusion declines smoothly and monotonically (illustrative magnitudes, lower is better)

  • Adversarial (GAN)
  • Denoising diffusion

仕組み

  1. 01

    Fix the distribution to model

    First decide the target: reproduce the whole dataset unconditionally, p(x), or generate under a condition such as text or a class label, p(x|y). This choice fixes the model’s output interface and where its training signal comes from.

  2. 02

    Choose how to represent the density

    Explicit-density models (autoregressive, VAE, diffusion) write down or approximate a functional form for p(x) and can be trained by likelihood; implicit-density models (GAN) never give a probability, only a sampler, and learn from an adversarial real-or-fake signal. The former has a clear objective but is constrained; the latter is flexible but hard to train.

  3. 03

    Optimise with likelihood or an adversarial signal

    Autoregressive models factorise the joint probability into a product of per-step predictions; VAEs optimise a lower bound on the data likelihood; diffusion splits generation into hundreds of denoising steps; GANs let a discriminator serve as a learned loss. Four routes, four mathematical formulations of "looking real".

  4. 04

    Measure quality with metrics and human judgement

    FID compares the generated and real distributions in Inception feature space, and IS scores the sharpness and diversity of individual samples — both are only proxies. The trustworthy judge remains human preference, whether by survey or pairwise arena voting.

図 3

Four principal routes: each gives a different mathematical definition of "looking real"

Generative modelsGenerative modelsAutoregressiveAutoregressivePer-step prediction, clear likelihoodPer-step prediction, clear l…AutoencoderAutoencoderCompress to a bottleneck, then rebuildCompress to a bottleneck, th…AdversarialAdversarialGenerator versus discriminator, implicit densityGenerator versus discriminat…DiffusionDiffusionStep-by-step denoising, stable to trainStep-by-step denoising, stab…
図 4

Representative FID magnitudes on ImageNet 256×256 by route (lower is better; autoregressive as VQGAN+Transformer, autoencoder as VQGAN, adversarial as BigGAN-deep, diffusion as DiT-XL/2)

応用場面

  • Image and video generation: synthesising new visual content from text, sketches or reference images
  • Speech and music synthesis: text-to-speech, singing-voice synthesis, music continuation
  • Data augmentation and simulation: synthesising samples for rare classes, or synthetic scenes for autonomous driving
  • Scientific design: generating candidate molecules, protein structures or chip layouts

よくある誤解

  • Realistic output does not imply the right distribution was learned. A model that is sharp but low in diversity — mode collapse — can look good on FID while covering only a small part of the data.
  • FID and IS depend on a fixed feature network; swap the extractor and the ranking can flip. They correlate with human intuition, they do not equal it.
  • There is no answer key for generative output. One prompt admits countless reasonable results, so evaluation must separate "implausible" from merely "unusual".

重要用語

Explicit density
A model that writes down or approximates p(x), e.g. autoregressive or diffusion
Implicit density
A model that offers only a sampler, not a probability, e.g. a GAN
Mode collapse
When a generator covers only a few modes of the data distribution
FID
Fréchet distance between generated and real distributions in Inception feature space; lower is better

参考文献