인간 피드백 기반 강화학습
‘좋은 답’을 수식으로 쓸 수 없을 때는 인간이 보상 함수 역할을 한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Reinforcement learning from human feedback (RLHF) targets tasks whose reward is hard to hand-write: it first trains a reward model on human preferences between candidate outputs, then uses that reward model as a differentiable proxy reward to fine-tune a generative model with RL (usually PPO). It is among the mainstream methods for aligning large language models with human preferences; the core idea is to replace scoring with comparing, which is cheaper and more consistent.
직관적 이해
Teaching a child to "write well" is hard to reduce to a precise formula, yet saying "this essay is better" between two is easy. RLHF collects this cheaper feedback, trains a model to imitate your preferences, and uses it as an automatic scorer. Humans only compare; the machine scores. The real risk lies exactly here: the machine learns "what you think you like", and once a bias creeps in, the model optimises straight along it.
The three-stage RLHF pipeline: supervised fine-tuning lays the base, human preferences train a reward model, and PPO optimises the policy under reward plus a KL penalty; the SFT model serves as both policy start and reference
Win rate in human-preference comparisons across alignment stages (magnitudes follow the published InstructGPT results: after RLHF, smaller models beat much larger unaligned ones in most comparisons)
작동 원리
- 01
Step 1 · Supervised fine-tuning as the base
Starting from the pretrained model, supervised-fine-tune on high-quality demonstrations to get a model that basically follows instructions. It is both the starting point of the RL stage and the floor on capability — a reward model can only steer, not fabricate ability.
- 02
Step 2 · Train the reward model
Show annotators two answers to the same prompt and have them pick the better one. Train a scoring model on these pairwise comparisons with a loss that rewards correct ordering, most often the Bradley–Terry form. Its output score is a differentiable approximation of human preference.
- 03
Step 3 · Optimise with PPO under a KL penalty
The language model acts as the policy, token-by-token generation is the action sequence, and the reward model’s score is the return. PPO maximises "reward − β·KL(new policy ‖ reference policy)". The KL term is essential: it keeps the model from drifting into degenerate language while trying to game the reward model.
- 04
Step 4 · A simpler substitute: DPO
DPO merges the two stages — train a reward model, then run RL — into a single closed-form loss built directly on preference pairs, removing sampling and the RL loop, which makes it simpler and steadier. The trade-off is losing on-policy exploration, so results depend heavily on the quality and coverage of the preference data.
핵심 수식
max_π E[ r_θ(x, y) ] − β · KL( π(y | x) ‖ π_ref(y | x) )응용 분야
- Chat assistant alignment: making responses more helpful, more honest and less harmful
- Summarisation and translation: tuning style, faithfulness and conciseness to human preference
- Code assistants: preference data steers generation toward code that matches developers’ habits
- Image and multimodal generation: similar preference tuning is used to match users’ aesthetic judgement
흔한 오해
- Reward hacking: the model learns the reward model’s preferences rather than the real ones. It may learn to flatter, to pad answers with length, or to fabricate plausible-sounding but unfounded responses to score highly.
- The reward model is itself limited: trained on finite preference data, its scores are unreliable out of distribution, on very long conversations and on complex reasoning, and upstream errors get amplified downstream.
- KL and β must be balanced: too weak and the model drifts into degenerate language; too strong and it barely learns, effectively freezing the reference policy. That balance point usually has to be re-tuned per task.
핵심 용어
- Reward model
- A model fitting human preferences and emitting a differentiable score
- Preference pair
- Two candidate outputs for one input plus the human’s choice between them
- KL penalty
- Penalises divergence from the reference policy to prevent degeneration
- Reward hacking
- Exploiting the proxy reward instead of genuinely completing the task