Aller au contenu
Atlas de l'IA

Fonctions de valeur et Q-learning

Plutôt que de deviner quoi faire, estimez d’abord la valeur de chaque choix

06 Apprentissage par renforcementIntermédiaireEntrée 2 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

The state value V(s) is the expected discounted return from state s under a given policy; the action value Q(s, a) additionally fixes the first action. Both satisfy the Bellman equation, which restates current value as "immediate reward + discounted value of the next state" — a self-consistent relation that permits bootstrapped, iterative approximation. Q-learning is the best-known instance: it approximates the optimal action value directly and can learn the optimal policy while still exploring.

Intuition

Think of each state as a cell holding "on average, how many points you can win from here". Q-learning repeatedly corrects the current cell using "the points from this step plus the estimated value of the next cell". The next cell may itself be wrong, but each pass makes them all a little truer at once. It is like calibrating several inaccurate clocks against one another: they converge together, provided no single clock is allowed to jump wildly.

Fig. 1

One Q-learning update: the target combines the immediate reward with the max Q of the next state; their difference is the TD error that corrects the current Q

Fig. 2

A Q-table learned in a simple maze: rows are states, columns are actions, and the brighter the cell the higher the action’s expected return; the exit row is flat because every action is "already arrived"

Fonctionnement

  1. 01

    The Bellman equation: value is self-consistent

    V(s) = E[ r + γ V(s′) ]. In principle this is a huge linear system with one unknown per state; solving it directly is infeasible, but the equation hands you the entry point for iteration.

  2. 02

    Temporal difference: correct this step with the next

    After each step, take "the reward actually received + γ times the old estimate of the next state" as the target, and update by its difference from the old estimate of the current state. That difference is the TD error, the single most important learning signal in reinforcement learning.

  3. 03

    Off-policy: always aiming at the optimum

    In Q-learning the update uses the maximum Q over all actions in the next state, not the action the agent would actually take. So it learns the value of the optimal policy while behaving randomly — behaviour is decoupled from learning, which is why it is called off-policy and why its sample efficiency is high.

  4. 04

    From tables to function approximation

    With few states, storing one number per (s, a) suffices; as states multiply — continuous, or raw pixels — the table explodes. The fix is a parameterised function Q(s, a; θ). The gradient descent looks almost like supervised learning, with one difference: the target itself moves along with the parameters.

Formule clé

Q(s, a) ← Q(s, a) + α [ r + γ · max_{a′} Q(s′, a′) − Q(s, a) ]
The Q-learning update. The bracket is the TD error, α the learning rate; taking the max rather than the actual action is precisely what makes it off-policy.

Où c'est utilisé

  • Classic control and games: from CartPole to Atari, Q-learning is the default starting point for discrete-action tasks
  • Recommendation and ranking: treat each impression as a step, and Q values yield a long-horizon serving policy
  • Robot navigation: shortest paths on a grid map can be approximated step by step with Q-learning
  • As a building block: the critic in actor-critic methods is usually a Q or V network

Idées fausses courantes

  • Overestimation bias: the max operation repeatedly picks the noisy side that happens to be higher, so Q-learning systematically overestimates true value. Double Q-learning corrects this by having two estimates check each other.
  • Using bootstrapping, off-policy learning and function approximation together is called the "deadly triad" and can diverge. It is exactly why deep Q-networks need experience replay and target networks.
  • Tabular methods work only when states can be enumerated. Swapping the index for a neural network does not by itself fix sample efficiency — at the same number of interactions, Q-learning can still lag far behind a human.

Termes clés

State value V(s)
Expected discounted return from s under policy π
Action value Q(s, a)
Expected discounted return after forcing the first action to be a
TD error
The gap between the fresh target and the old estimate
Off-policy
The behaviour policy may differ from the policy being learned

Lectures complémentaires