強化学習
試行錯誤と遅延報酬から一連の意思決定を学ぶ
OVERVIEW
- 全項目
- 6
- 初級
- 2
- 中級
- 3
- 上級
- 1
Reinforcement learning addresses problems that have no answer key, only consequences: whether a step was good may only be revealed much later as reward. It introduces the vocabulary of agent, environment, state, action and reward, and approaches optimal behaviour through two families of ideas — value functions and policy gradients.
この領域が答える問い
- Q1
Why does delayed reward make the problem hard?
- Q2
How do value functions and policy gradients divide the work?
- Q3
Why can RL improve large language models?