Aller au contenu
Atlas de l'IA

Apprentissage par renforcement profond

Laissez un réseau de neurones décider directement depuis les pixels, stabilisé par des idées anciennes

06 Apprentissage par renforcementIntermédiaireEntrée 5 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

Deep reinforcement learning uses deep networks to approximate value functions or policies, removing the tabular method’s limit on the number of states so agents can learn control directly from high-dimensional inputs such as image pixels. It couples RL’s bootstrapped targets with neural function approximation, but these two interfere with each other, so techniques such as experience replay and target networks are required to train stably.

Intuition

Everything before assumed states could be listed in a table. But from a single 210×160 game frame there are more possible images than atoms in the universe — the table can never be written. Deep RL’s idea is direct: replace the table with a network that looks at images. Trouble follows: the network changes while the target it chases changes too, like aiming at a moving target while running. So the target is pinned down for a while (target network) and past experience is stored and replayed (experience replay).

Fig. 1

The AlphaZero self-play loop: the network guides search, search produces stronger games, and those games train a stronger network — capability climbs on its own inside the loop

Network guides MCTSGenerate games and…Update networks fr…Stronger network →…Self-play
Fig. 2

Human-normalised scores of DQN on selected Atari games (100% = human level): sharply bipolar — far above human on Boxing and Beam Rider, yet almost zero on Montezuma’s Revenge

Fonctionnement

  1. 01

    Replay and target networks: two stabilisers

    Store each (s, a, r, s′) step in a replay buffer and sample random mini-batches during training. This breaks the temporal correlation between samples and lets one experience be reused many times. A separate frozen target network computes the next-state value, so the target does not move around together with the parameters.

  2. 02

    Atari: from pixels to human level

    In 2015 DQN used one architecture and one hyperparameter set on 49 Atari games, working only from raw pixels and the game score, to match or beat all previous specialised methods on most of them and to surpass human play on roughly half. It was the first demonstration that a single model can span many tasks.

  3. 03

    AlphaGo / AlphaZero: search plus self-play

    AlphaGo combined policy and value networks with Monte Carlo tree search: the networks suggest where to look and the search does the careful calculation. AlphaZero went further, dropping human games and learning purely from self-play under one general rule set, surpassing the top programs in Go, chess and shogi.

  4. 04

    Sample efficiency: the biggest bottleneck

    DQN needed roughly 50 million frames to master one Atari game, while a human needs tens of minutes. Counting every interaction with a simulator or real hardware as costly, deep RL is extremely expensive on physical systems. This is why model-based, offline RL and imitation learning receive sustained attention as alternatives.

Où c'est utilisé

  • Game AI: from Atari to StarCraft II and Dota 2, top systems rely heavily on deep RL
  • Robotic manipulation: grasping, locomotion and dexterous hands, usually trained in simulation and transferred to hardware
  • Industrial and scientific control: data-centre cooling, chip placement, tokamak plasma control
  • LLM alignment: RLHF is, at bottom, deep RL applied to a language model

Idées fausses courantes

  • The "deadly triad": bootstrapping, off-policy learning and function approximation can diverge when combined. Target networks, replay and gradient clipping suppress it in practice, but do not eliminate it.
  • Extremely sensitive to hyperparameters and seeds. The same algorithm with a different seed can differ several-fold in final performance, making reproduction and fair comparison notoriously hard.
  • Good performance in simulation does not imply transfer to the real world. The sim-to-real gap appears in perception, dynamics and latency alike, and usually calls for extra measures such as domain randomisation.

Termes clés

Experience replay
Store past transitions and sample randomly to break correlation
Target network
A slowly updated copy of the network providing stable bootstrap targets
MCTS
An algorithm that evaluates moves via sampled rollouts to guide search
Self-play
Generating training data by having an agent play against its past selves

Lectures complémentaires