मुख्य सामग्री पर जाएँ
01 गणित और सांख्यिकी की मूल बातेंप्रारंभिकइस क्षेत्र की प्रविष्टि 2

प्रवणता और प्रवणता अवरोहण

संपूर्ण गहन शिक्षण एक ही बात पर टिका है: ढलान की ओर एक छोटा कदम

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

परिभाषा

The gradient is a vector pointing in the direction of steepest increase; its opposite is the direction of steepest decrease. Gradient descent is an iterative algorithm: compute the gradient of the loss at the current parameters, move the parameters a small distance in the opposite direction, and repeat until the loss stops improving.

सहज समझ

Picture yourself on a hillside in thick fog: you cannot see the whole terrain, only which way the ground falls away beneath your feet. The pragmatic strategy is not to leap to the valley floor but to take one small step in the steepest downward direction, then feel the ground again. Steps too large will overshoot the valley — possibly climbing higher; steps too small will never arrive before dark. The "learning rate" is simply how far you step.

चित्र 1

The geometry of gradient descent: each step follows the steepest descent from where you stand, scaled by the learning rate

चित्र 2

The learning rate decides success: three runs share identical data and initial weights, differing only in step size (click the legend to toggle a curve)

  • Too large (oscillates, diverges)
  • Appropriate (converges steadily)
  • Too small (barely moves)

कार्यप्रणाली

  1. 01

    Forward: measure how wrong you currently are

    Push a batch of samples through the model, obtain predictions, compare them with the truth, and reduce the difference to a single scalar loss. That scalar is the elevation of the hill you are standing on.

  2. 02

    Backward: automatic differentiation gives the direction

    Backpropagation walks the computation graph from output back to input, applying the chain rule layer by layer, obtaining in one pass the partial derivative of the loss with respect to every parameter. This is the central trick of modern frameworks: turning "derive the gradient" from mathematical drudgery into a single function call.

  3. 03

    Update: take the step

    Parameter ← parameter − learning rate × gradient. One step looks trivial, yet after millions of them a randomly initialised network grows internal structure that can recognise speech and write code.

  4. 04

    Momentum and adaptivity: making the step smarter

    Plain gradient descent zig-zags inside a narrow valley. Momentum accumulates past directions to smooth the oscillation, while Adam tunes the step size per parameter — the default optimiser for essentially every large model today.

मुख्य सूत्र

θ ← θ − η · ∇θ L(θ)
θ is the parameter vector, η the learning rate and ∇θL the gradient of the loss. This is the skeleton of every deep-learning training loop.
चित्र 4

The typical character of three optimisers: convergence speed, stability, sensitivity to hyperparameters and memory cost

  • SGD + momentum
  • Adam
  • RMSProp

उपयोग के क्षेत्र

  • Training every neural network: from handwritten-digit recognition to hundred-billion-parameter language models
  • Fine-tuning: continue descending on a new task with a tiny learning rate so existing abilities are not destroyed
  • Reinforcement learning: treat policy performance as the loss, so gradient ascent becomes policy improvement
  • Adversarial examples: perturbing the input along the gradient can flip a model’s judgement entirely

सामान्य भ्रांतियाँ

  • A gradient is purely local. It tells you the steepest direction underfoot, not whether that direction leads to the global optimum. Deep losses teem with local valleys and saddle points, yet in practice sufficiently wide networks are full of "good enough" local solutions.
  • The learning rate is the hyperparameter that matters most. On identical data, a tenfold difference can separate "converges beautifully" from "loss becomes NaN".
  • Vanishing and exploding gradients: the chain rule multiplies terms, so gradients in deep networks shrink or blow up exponentially — which is precisely why residual connections, normalisation and gradient clipping exist.

मुख्य शब्द

Learning rate η
How far each step moves
Batch size
How many samples estimate the gradient per step
SGD
Approximating the full gradient with a mini-batch
Loss surface
The high-dimensional terrain of loss values over parameter space

आगे का पठन