Gradient chính sách
Điều chỉnh trực tiếp chính sách để hành động tốt xuất hiện thường xuyên hơn
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
ĐỊNH NGHĨA
Policy-gradient methods write the policy as a parameterised distribution πθ(a | s) and take the gradient of expected return directly with respect to θ, ascending it so that high-return actions are sampled more often. Because they never estimate a value and then take a maximum, they naturally support continuous actions and stochastic policies — at the cost of high estimation variance.
Trực giác
Rather than scoring every state-action pair and picking the largest, ask directly: within the current policy, which action’s probability should go up and which down? The recipe is plain — run a few episodes and nudge up the probability of every action in a high-return trajectory, and down those in poor ones. The hard part is that "this episode scored high" is mixed with a great deal of luck, so many samples are needed to average the noise away and see the true direction.
Two routes: learn "how much it is worth" then take the max, or learn "how to act" directly — one is sample-efficient but awkward with continuous actions, the other flexible but high-variance
Learning curves on CartPole-v1 (maximum 500): PPO converges fastest and steadiest, while plain REINFORCE is slow and noisy — exactly what baselines, critics and trust regions address
- PPO (actor–critic + clipping)
- A2C (actor–critic with advantage)
- REINFORCE (with baseline)
Cách hoạt động
- 01
Log-derivative trick: make sampling differentiable
The central identity is ∇J = E[ ∇log πθ(a | s) · G ]. It rewrites "we cannot differentiate through sampling" into a weighted sum computable per sampled action — the larger the return G, the harder that action’s log-probability is pushed up.
- 02
REINFORCE and baselines: cut the variance first
The earliest algorithm used the whole-episode return G as the weight, hence the name REINFORCE, but its variance is huge. Subtracting a baseline b(s) that does not depend on the action — usually the state value V(s) — cuts the variance sharply without adding bias. This is the seed of the "advantage".
- 03
Actor–critic: one acts, one judges
Weight updates by A(s, a) = Q(s, a) − V(s) — "how much better than average this action is". A critic estimates it while an actor updates on it, trained in alternation. Compared with plain REINFORCE it replaces the whole-episode return with a per-step advantage, cutting variance.
- 04
Trust regions: TRPO and PPO
Step too far and the new policy collapses, wasting all previously collected data. TRPO explicitly constrains the KL divergence between old and new policies; PPO replaces it with a simpler clipped objective keeping the probability ratio inside [1−ε, 1+ε]. Easier to implement and steadier, PPO became today’s default algorithm.
Công thức then chốt
∇θ J ≈ E[ ∇θ log πθ(a | s) · A(s, a) ]The character of three routes: Q-learning is sample-efficient but weak on continuous control, pure policy gradients are flexible but high-variance, and actor–critic / PPO strikes a balance
- Q-learning
- REINFORCE
- PPO
Ứng dụng
- Continuous control: robotic arms, legged robots, steering and throttle for autonomous driving
- LLM alignment: the final stage of RLHF optimises the language model policy with PPO
- Stochastic games: settings that require a mixed strategy, such as rock-paper-scissors-style play
- Combinatorial optimisation and scheduling: the policy emits a distribution over actions, naturally handling random choices
Hiểu lầm thường gặp
- High variance: the return of a single trajectory fluctuates wildly, so gradient estimates are far noisier than in supervised learning and need many samples or a better advantage estimate to stabilise.
- The cost of being on-policy: once the policy updates, old data is stale and must be discarded. This makes sample efficiency inherently lower than off-policy Q-learning.
- Sensitivity to step size: too large an update can collapse the policy beyond recovery. That is precisely why trust regions and clipping exist, yet ε and the learning rate still need careful tuning.
Thuật ngữ chính
- Actor / Critic
- The policy network and the value network: one acts, one scores
- Advantage A(s, a)
- How much better an action is than the average at that state
- Log-derivative trick
- Turns the gradient of an expectation into a weighted sum of log-probabilities
- Trust region / KL constraint
- Bounds how far the new policy may drift from the old