Aller au contenu
Atlas de l'IA

Rétropropagation

Transformer « calculer le gradient » d’un pensum mathématique en un simple appel de fonction : l’instant où l’apprentissage profond a décollé

03 Apprentissage profondIntermédiaireEntrée 2 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

Backpropagation is an efficient algorithm for computing gradients. It first runs a forward pass over the computation graph to obtain the loss, then walks backward from the output to the inputs, applying the chain rule to derive the partial derivative of the loss with respect to every layer’s parameters. With it, a single backward sweep yields all gradients, no matter how many layers or parameters the network has.

Intuition

Picture a factory line: a product (the input) passes through stations (the layers) into a finished good (the prediction), and inspection measures the gap to spec (the loss). To trace how much each station should change, you walk back from the last station, distributing blame in proportion to contribution. Each station only needs the error returning from downstream, multiplies it by its own local derivative, and passes it further up — with no need to recompute the whole line. Backpropagation does exactly this, applying the chain rule precisely at every step.

Fig. 1

The two sweeps of one training iteration: the solid path is forward (computing the loss), the dashed path is backward (propagating gradients by the chain rule)

Fig. 2

Gradients decaying with depth: in an 8-layer network, Sigmoid makes layers near the input receive exponentially smaller gradients, whereas ReLU with normalised activations keeps them roughly even (log scale on the y-axis)

  • Sigmoid (deep decay)
  • ReLU + normalisation (stable)

Fonctionnement

  1. 01

    Forward pass: obtain the loss and cache intermediates

    Run the layers in order — a linear map followed by a nonlinearity — to obtain the prediction and a scalar loss, caching each layer’s intermediate activations for the backward pass.

  2. 02

    Backward pass: the chain rule, layer by layer

    Starting from the loss, multiply the gradient with respect to a layer’s output by that layer’s local derivative to obtain the gradient with respect to its input, then hand it to the previous layer. Each layer performs just one local multiplication.

  3. 03

    Obtain parameter gradients: hand them to the optimiser

    An outer product of the cached input with the incoming gradient (a convolution in convolutional networks) yields the gradients with respect to weights and biases, which a gradient-based optimiser consumes directly.

  4. 04

    The meaning of autodiff: no more structural limits

    A framework records any computation as a graph and performs backpropagation automatically. This is what freed researchers from hand-deriving partial derivatives and let them design arbitrarily deep, arbitrarily structured networks.

Formule clé

∂L/∂W⁽ˡ⁾ = δ⁽ˡ⁾ (a⁽ˡ⁻¹⁾)ᵀ, δ⁽ˡ⁾ = ((W⁽ˡ⁺¹⁾)ᵀ δ⁽ˡ⁺¹⁾) ⊙ f′(z⁽ˡ⁾)
The error δ of layer l is the next layer’s error propagated through the transposed weights, multiplied element-wise by this layer’s activation derivative.

Où c'est utilisé

  • The shared engine that trains every neural network — CNNs, RNNs and Transformers alike
  • Bound to autodiff libraries (PyTorch, TensorFlow, JAX) to form the standard training loop
  • Enables fine-tuning and parameter-efficient methods such as LoRA that keep backpropagating through existing weights
  • Analysing gradient flow: diagnosing vanishing/exploding gradients or drawing saliency maps to explain predictions

Idées fausses courantes

  • Backpropagation only computes gradients; it is distinct from the optimiser (SGD, Adam). The former yields a direction, the latter decides how to use it.
  • Backpropagation must retain the forward activations, so memory grows with depth and batch size; gradient checkpointing trades recomputation for less storage.
  • Gradients are local information. They give the steepest direction underfoot without any guarantee that it leads to the global optimum.

Termes clés

Computation graph
A computation expressed as nodes and directed edges over which derivatives propagate
Chain rule
The derivative of a composition is the product of the local derivatives
Automatic differentiation
Letting a framework compute exact gradients automatically, not by numerical approximation
Gradient flow
The magnitude and stability of gradients as they propagate layer by layer

Lectures complémentaires