Reinforcement Learning
Learning sequences of decisions from trial, error and delayed reward
OVERVIEW
- All Concepts
- 6
- Beginner
- 2
- Intermediate
- 3
- Expert
- 1
Reinforcement learning addresses problems that have no answer key, only consequences: whether a step was good may only be revealed much later as reward. It introduces the vocabulary of agent, environment, state, action and reward, and approaches optimal behaviour through two families of ideas — value functions and policy gradients.
Questions this domain answers
- Q1
Why does delayed reward make the problem hard?
- Q2
How do value functions and policy gradients divide the work?
- Q3
Why can RL improve large language models?
Concepts in this domain
- 01Markov Decision ProcessBeginnerWrite "deciding step by step" as five symbols — everything in reinforcement learning starts here
- 02Value Functions & Q-LearningIntermediateInstead of guessing what to do, first estimate what each choice is worth
- 03Policy GradientsIntermediateAdjust the policy itself, so that good actions occur more often
- 04Exploration vs ExploitationBeginnerThe best option right now is not necessarily the best one in the long run
- 05Deep Reinforcement LearningIntermediateLet a neural network decide straight from pixels — then hold it steady with decades-old tricks
- 06Reinforcement Learning from Human FeedbackExpertWhen the good answer cannot be written as a formula, let humans stand in as the reward function