Chuyển đến nội dung
Bản đồ AI

Gradient chính sách

Điều chỉnh trực tiếp chính sách để hành động tốt xuất hiện thường xuyên hơn

06 Học tăng cườngTrung cấpMục từ thứ 3 trong lĩnh vực

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

Policy-gradient methods write the policy as a parameterised distribution πθ(a | s) and take the gradient of expected return directly with respect to θ, ascending it so that high-return actions are sampled more often. Because they never estimate a value and then take a maximum, they naturally support continuous actions and stochastic policies — at the cost of high estimation variance.

Trực giác

Rather than scoring every state-action pair and picking the largest, ask directly: within the current policy, which action’s probability should go up and which down? The recipe is plain — run a few episodes and nudge up the probability of every action in a high-return trajectory, and down those in poor ones. The hard part is that "this episode scored high" is mixed with a great deal of luck, so many samples are needed to average the noise away and see the true direction.

Hình 1

Two routes: learn "how much it is worth" then take the max, or learn "how to act" directly — one is sample-efficient but awkward with continuous actions, the other flexible but high-variance

Hình 2

Learning curves on CartPole-v1 (maximum 500): PPO converges fastest and steadiest, while plain REINFORCE is slow and noisy — exactly what baselines, critics and trust regions address

  • PPO (actor–critic + clipping)
  • A2C (actor–critic with advantage)
  • REINFORCE (with baseline)

Cách hoạt động

  1. 01

    Log-derivative trick: make sampling differentiable

    The central identity is ∇J = E[ ∇log πθ(a | s) · G ]. It rewrites "we cannot differentiate through sampling" into a weighted sum computable per sampled action — the larger the return G, the harder that action’s log-probability is pushed up.

  2. 02

    REINFORCE and baselines: cut the variance first

    The earliest algorithm used the whole-episode return G as the weight, hence the name REINFORCE, but its variance is huge. Subtracting a baseline b(s) that does not depend on the action — usually the state value V(s) — cuts the variance sharply without adding bias. This is the seed of the "advantage".

  3. 03

    Actor–critic: one acts, one judges

    Weight updates by A(s, a) = Q(s, a) − V(s) — "how much better than average this action is". A critic estimates it while an actor updates on it, trained in alternation. Compared with plain REINFORCE it replaces the whole-episode return with a per-step advantage, cutting variance.

  4. 04

    Trust regions: TRPO and PPO

    Step too far and the new policy collapses, wasting all previously collected data. TRPO explicitly constrains the KL divergence between old and new policies; PPO replaces it with a simpler clipped objective keeping the probability ratio inside [1−ε, 1+ε]. Easier to implement and steadier, PPO became today’s default algorithm.

Công thức then chốt

∇θ J ≈ E[ ∇θ log πθ(a | s) · A(s, a) ]
The policy gradient with an advantage. Replacing the full return G with A(s, a) is the key trade-off between variance and bias.
Hình 4

The character of three routes: Q-learning is sample-efficient but weak on continuous control, pure policy gradients are flexible but high-variance, and actor–critic / PPO strikes a balance

  • Q-learning
  • REINFORCE
  • PPO

Ứng dụng

  • Continuous control: robotic arms, legged robots, steering and throttle for autonomous driving
  • LLM alignment: the final stage of RLHF optimises the language model policy with PPO
  • Stochastic games: settings that require a mixed strategy, such as rock-paper-scissors-style play
  • Combinatorial optimisation and scheduling: the policy emits a distribution over actions, naturally handling random choices

Hiểu lầm thường gặp

  • High variance: the return of a single trajectory fluctuates wildly, so gradient estimates are far noisier than in supervised learning and need many samples or a better advantage estimate to stabilise.
  • The cost of being on-policy: once the policy updates, old data is stale and must be discarded. This makes sample efficiency inherently lower than off-policy Q-learning.
  • Sensitivity to step size: too large an update can collapse the policy beyond recovery. That is precisely why trust regions and clipping exist, yet ε and the learning rate still need careful tuning.

Thuật ngữ chính

Actor / Critic
The policy network and the value network: one acts, one scores
Advantage A(s, a)
How much better an action is than the average at that state
Log-derivative trick
Turns the gradient of an expectation into a weighted sum of log-probabilities
Trust region / KL constraint
Bounds how far the new policy may drift from the old

Đọc thêm