Aller au contenu
Atlas de l'IA

Réseaux de neurones récurrents

Donner une mémoire au réseau : une même unité réutilisée dans le temps pour traiter des séquences de longueur quelconque

03 Apprentissage profondIntermédiaireEntrée 5 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

A recurrent neural network processes sequences. At each time step it reads an input and receives the previous hidden state, producing a new hidden state and an output. Because the same unit is reused over time and all steps share one set of weights, it handles variable-length sequences and carries history into the current decision.

Intuition

Reading a sentence is like listening to someone speak: you do not replay the whole thing, you keep updating an in-head summary of "what has been said so far" — that continuously rewritten summary is the hidden state. Sadly, a vanilla RNN’s memory is short: information is multiplied at every step, so after dozens of steps earlier content either fades (vanishing gradient) or blows up (exploding gradient). Gating is, at heart, the act of fitting that summary with a few valves: what to store, what to drop, what to reveal.

Fig. 1

A recurrent network unrolled over time: the same unit reads xₜ, updates hₜ and emits yₜ at each step, with hidden states linked by one shared weight Wₕ

Fig. 2

The longer the sequence, the harder it is for a vanilla RNN to retain early information, and accuracy collapses; gated LSTM/GRU stay far more stable on long sequences (a long-range-dependency benchmark; magnitudes from typical reported results)

  • Vanilla RNN
  • LSTM
  • GRU

Fonctionnement

  1. 01

    Hidden state and unrolling over time

    hₜ = f(W hₜ₋₁ + U xₜ + b). Although called "recurrent", training unrolls it into a very deep chain over time, with all steps sharing a single set of weights.

  2. 02

    Vanishing and exploding gradients

    Errors are propagated back through time (BPTT), multiplied by the weight matrix and activation derivative at each step. This repeated product makes the gradient decay or grow exponentially with the number of steps, so long-range dependencies are hard to learn.

  3. 03

    Gating: LSTM and GRU

    LSTM uses input, forget and output gates to decide what to store, drop and reveal, protecting gradients through a near-identity cell-state path; GRU achieves a similar effect with fewer gates.

  4. 04

    Superseded by attention

    Even with gating, an RNN must compute step by step, cannot parallelise like a Transformer, and struggles to retain very long context. Attention replaces "pass it along word by word" with "look at all positions directly", which is why it became the mainstream.

Où c'est utilisé

  • Language modelling and machine translation (the mainstream before attention became widespread)
  • Frame-by-frame tasks such as speech and handwriting recognition
  • Time-series forecasting: sensor, weather and financial data
  • Encoders and decoders for sequence-to-sequence tasks

Idées fausses courantes

  • "RNNs have memory" only to a degree: a vanilla RNN’s effective memory is typically a few dozen steps, and information beyond that is hard to retain.
  • LSTM mitigates rather than solves long-range dependencies. Gating adds computation without changing the fundamental sequential bottleneck.
  • Training with teacher forcing but generating freely at inference creates a distribution mismatch that accumulates and amplifies errors.

Termes clés

Hidden state
A continuously updated "summary so far" vector
BPTT
Backpropagation through time after unrolling
Gating
Using 0–1 coefficients from Sigmoid to control how much information passes
Long-range dependency
Influence between elements far apart in a sequence

Lectures complémentaires