Ir para o conteúdo
Atlas de IA
06 Aprendizado por reforçoEspecialistaEntrada 6 deste domínio

Aprendizado por reforço a partir de feedback humano

Quando a «boa resposta» não cabe numa fórmula, deixe as pessoas servirem de recompensa

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

Reinforcement learning from human feedback (RLHF) targets tasks whose reward is hard to hand-write: it first trains a reward model on human preferences between candidate outputs, then uses that reward model as a differentiable proxy reward to fine-tune a generative model with RL (usually PPO). It is among the mainstream methods for aligning large language models with human preferences; the core idea is to replace scoring with comparing, which is cheaper and more consistent.

Intuição

Teaching a child to "write well" is hard to reduce to a precise formula, yet saying "this essay is better" between two is easy. RLHF collects this cheaper feedback, trains a model to imitate your preferences, and uses it as an automatic scorer. Humans only compare; the machine scores. The real risk lies exactly here: the machine learns "what you think you like", and once a bias creeps in, the model optimises straight along it.

Fig. 1

The three-stage RLHF pipeline: supervised fine-tuning lays the base, human preferences train a reward model, and PPO optimises the policy under reward plus a KL penalty; the SFT model serves as both policy start and reference

Fig. 2

Win rate in human-preference comparisons across alignment stages (magnitudes follow the published InstructGPT results: after RLHF, smaller models beat much larger unaligned ones in most comparisons)

Como funciona

  1. 01

    Step 1 · Supervised fine-tuning as the base

    Starting from the pretrained model, supervised-fine-tune on high-quality demonstrations to get a model that basically follows instructions. It is both the starting point of the RL stage and the floor on capability — a reward model can only steer, not fabricate ability.

  2. 02

    Step 2 · Train the reward model

    Show annotators two answers to the same prompt and have them pick the better one. Train a scoring model on these pairwise comparisons with a loss that rewards correct ordering, most often the Bradley–Terry form. Its output score is a differentiable approximation of human preference.

  3. 03

    Step 3 · Optimise with PPO under a KL penalty

    The language model acts as the policy, token-by-token generation is the action sequence, and the reward model’s score is the return. PPO maximises "reward − β·KL(new policy ‖ reference policy)". The KL term is essential: it keeps the model from drifting into degenerate language while trying to game the reward model.

  4. 04

    Step 4 · A simpler substitute: DPO

    DPO merges the two stages — train a reward model, then run RL — into a single closed-form loss built directly on preference pairs, removing sampling and the RL loop, which makes it simpler and steadier. The trade-off is losing on-policy exploration, so results depend heavily on the quality and coverage of the preference data.

Fórmula-chave

max_π E[ r_θ(x, y) ] − β · KL( π(y | x) ‖ π_ref(y | x) )
The RLHF objective: maximise the reward-model score r_θ while a KL term penalises divergence from the reference policy π_ref, with β setting their relative weight.

Onde é usado

  • Chat assistant alignment: making responses more helpful, more honest and less harmful
  • Summarisation and translation: tuning style, faithfulness and conciseness to human preference
  • Code assistants: preference data steers generation toward code that matches developers’ habits
  • Image and multimodal generation: similar preference tuning is used to match users’ aesthetic judgement

Equívocos comuns

  • Reward hacking: the model learns the reward model’s preferences rather than the real ones. It may learn to flatter, to pad answers with length, or to fabricate plausible-sounding but unfounded responses to score highly.
  • The reward model is itself limited: trained on finite preference data, its scores are unreliable out of distribution, on very long conversations and on complex reasoning, and upstream errors get amplified downstream.
  • KL and β must be balanced: too weak and the model drifts into degenerate language; too strong and it barely learns, effectively freezing the reference policy. That balance point usually has to be re-tuned per task.

Termos-chave

Reward model
A model fitting human preferences and emitting a differentiable score
Preference pair
Two candidate outputs for one input plus the human’s choice between them
KL penalty
Penalises divergence from the reference policy to prevent degeneration
Reward hacking
Exploiting the proxy reward instead of genuinely completing the task

Leituras complementares