मुख्य सामग्री पर जाएँ

नीति प्रवणता

नीति को सीधे समायोजित करें ताकि अच्छे कार्य अधिक बार हों

06 प्रबलन अधिगममध्यवर्तीइस क्षेत्र की प्रविष्टि 3

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

परिभाषा

Policy-gradient methods write the policy as a parameterised distribution πθ(a | s) and take the gradient of expected return directly with respect to θ, ascending it so that high-return actions are sampled more often. Because they never estimate a value and then take a maximum, they naturally support continuous actions and stochastic policies — at the cost of high estimation variance.

सहज समझ

Rather than scoring every state-action pair and picking the largest, ask directly: within the current policy, which action’s probability should go up and which down? The recipe is plain — run a few episodes and nudge up the probability of every action in a high-return trajectory, and down those in poor ones. The hard part is that "this episode scored high" is mixed with a great deal of luck, so many samples are needed to average the noise away and see the true direction.

चित्र 1

Two routes: learn "how much it is worth" then take the max, or learn "how to act" directly — one is sample-efficient but awkward with continuous actions, the other flexible but high-variance

चित्र 2

Learning curves on CartPole-v1 (maximum 500): PPO converges fastest and steadiest, while plain REINFORCE is slow and noisy — exactly what baselines, critics and trust regions address

  • PPO (actor–critic + clipping)
  • A2C (actor–critic with advantage)
  • REINFORCE (with baseline)

कार्यप्रणाली

  1. 01

    Log-derivative trick: make sampling differentiable

    The central identity is ∇J = E[ ∇log πθ(a | s) · G ]. It rewrites "we cannot differentiate through sampling" into a weighted sum computable per sampled action — the larger the return G, the harder that action’s log-probability is pushed up.

  2. 02

    REINFORCE and baselines: cut the variance first

    The earliest algorithm used the whole-episode return G as the weight, hence the name REINFORCE, but its variance is huge. Subtracting a baseline b(s) that does not depend on the action — usually the state value V(s) — cuts the variance sharply without adding bias. This is the seed of the "advantage".

  3. 03

    Actor–critic: one acts, one judges

    Weight updates by A(s, a) = Q(s, a) − V(s) — "how much better than average this action is". A critic estimates it while an actor updates on it, trained in alternation. Compared with plain REINFORCE it replaces the whole-episode return with a per-step advantage, cutting variance.

  4. 04

    Trust regions: TRPO and PPO

    Step too far and the new policy collapses, wasting all previously collected data. TRPO explicitly constrains the KL divergence between old and new policies; PPO replaces it with a simpler clipped objective keeping the probability ratio inside [1−ε, 1+ε]. Easier to implement and steadier, PPO became today’s default algorithm.

मुख्य सूत्र

∇θ J ≈ E[ ∇θ log πθ(a | s) · A(s, a) ]
The policy gradient with an advantage. Replacing the full return G with A(s, a) is the key trade-off between variance and bias.
चित्र 4

The character of three routes: Q-learning is sample-efficient but weak on continuous control, pure policy gradients are flexible but high-variance, and actor–critic / PPO strikes a balance

  • Q-learning
  • REINFORCE
  • PPO

उपयोग के क्षेत्र

  • Continuous control: robotic arms, legged robots, steering and throttle for autonomous driving
  • LLM alignment: the final stage of RLHF optimises the language model policy with PPO
  • Stochastic games: settings that require a mixed strategy, such as rock-paper-scissors-style play
  • Combinatorial optimisation and scheduling: the policy emits a distribution over actions, naturally handling random choices

सामान्य भ्रांतियाँ

  • High variance: the return of a single trajectory fluctuates wildly, so gradient estimates are far noisier than in supervised learning and need many samples or a better advantage estimate to stabilise.
  • The cost of being on-policy: once the policy updates, old data is stale and must be discarded. This makes sample efficiency inherently lower than off-policy Q-learning.
  • Sensitivity to step size: too large an update can collapse the policy beyond recovery. That is precisely why trust regions and clipping exist, yet ε and the learning rate still need careful tuning.

मुख्य शब्द

Actor / Critic
The policy network and the value network: one acts, one scores
Advantage A(s, a)
How much better an action is than the average at that state
Log-derivative trick
Turns the gradient of an expectation into a weighted sum of log-probabilities
Trust region / KL constraint
Bounds how far the new policy may drift from the old

आगे का पठन