Ir para o conteúdo
Atlas de IA

Redes generativas adversárias

Um falsificador contra um perito: cada um leva o outro ao limite

07 IA generativaInicianteEntrada 3 deste domínio

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

A generative adversarial network is two networks: a generator G that maps random noise into samples, and a discriminator D that decides whether a sample came from real data or from G. The two improve through competition — G wants D to be fooled, D wants to tell them apart — and training ideally converges when the generated and real distributions coincide.

Intuição

This is a sustained two-player game. G is the counterfeiter, D the bank inspector. Each batch the counterfeiter produces, the inspector points out the flaw; each time the inspector gets sharper, the counterfeiter must render finer detail. When the fakes are indistinguishable even to the inspector, the game reaches equilibrium — GAN’s only, and fragile, objective. The danger: if one player learns much faster it crushes the other, and the game can oscillate or collapse.

Fig. 1

The adversarial loop: noise becomes a sample via G, D issues a real/fake verdict, whose signal flows back to G

Noise zGenerator GFake sampleDiscriminator DReal / fake verdictAdversarial…
Fig. 2

In a healthy GAN, discriminator and generator losses chase each other around log 2 ≈ 0.69 rather than monotonically decreasing; such oscillation is the normal signature of convergence to equilibrium

  • Discriminator loss
  • Generator loss

Como funciona

  1. 01

    Sample and generate

    Draw a noise vector from a simple prior — say a standard normal or uniform — and push it through the generator to obtain a synthetic sample. The generator does not copy any training image; it learns to transport the noise distribution onto the data distribution.

  2. 02

    Discriminate and score

    The discriminator receives both real and generated samples and outputs the probability that they are real. Its penalty for misjudging them becomes the training signal for the generator — effectively a learned loss function that keeps sharpening as the game proceeds.

  3. 03

    The minimax game

    The objective is a minimax problem: D maximises the log-likelihood of judging correctly while G minimises its chance of being detected. With an optimal discriminator, G’s gradient is equivalent to minimising the Jensen–Shannon divergence between the generated and real distributions.

  4. 04

    Improvements for stability

    DCGAN stabilises training with convolutional architectures and batch normalisation; WGAN adopts the Wasserstein distance with weight clipping or a gradient penalty, keeping gradients meaningful when distributions barely overlap; conditional GANs concatenate a label or text into both networks’ inputs to enable controlled generation.

Fórmula-chave

min_G max_D E_x[log D(x)] + E_z[log(1 − D(G(z)))]
The minimax objective: D maximises correct real/fake judgements, G minimises its chance of detection. At the ideal equilibrium the generated distribution equals the data distribution.
Fig. 3

GAN versus diffusion: the price of one-step generation is a highly unstable adversarial optimisation

Onde é usado

  • Image synthesis, super-resolution and image-to-image translation (e.g. CycleGAN)
  • Style transfer and restoration of old photographs
  • Data augmentation: synthesising samples for rare classes
  • As a component in other models: perceptual losses, and discriminators inside speech and video synthesis

Equívocos comuns

  • Mode collapse: G discovers that a few samples fool D best and keeps producing them, so diversity collapses.
  • Unstable training: G and D must progress in step; if one dominates, gradients vanish or oscillate, and hyperparameters are extremely sensitive.
  • A discriminator accuracy of 50% does not mean success. A D that is too weak or too strong both leave G learning nothing, and the balance is narrow.

Termos-chave

Generator
The network mapping noise to samples
Discriminator
The network judging real versus fake, serving as the loss
Mode collapse
The generator covers few modes and loses diversity
Wasserstein distance
An earth-mover distance between distributions, better behaved for training than JS divergence

Leituras complementares