본문으로 건너뛰기
AI 도감

강화학습

시행착오와 지연 보상으로 일련의 결정을 학습한다

OVERVIEW

전체 항목
6
입문
2
중급
3
고급
1

Reinforcement learning addresses problems that have no answer key, only consequences: whether a step was good may only be revealed much later as reward. It introduces the vocabulary of agent, environment, state, action and reward, and approaches optimal behaviour through two families of ideas — value functions and policy gradients.