التدرّج والانحدار التدريجي
يتلخّص التعلّم العميق كله في أمر واحد: خطوة صغيرة نحو الأسفل
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
The gradient is a vector pointing in the direction of steepest increase; its opposite is the direction of steepest decrease. Gradient descent is an iterative algorithm: compute the gradient of the loss at the current parameters, move the parameters a small distance in the opposite direction, and repeat until the loss stops improving.
الحدس المباشر
Picture yourself on a hillside in thick fog: you cannot see the whole terrain, only which way the ground falls away beneath your feet. The pragmatic strategy is not to leap to the valley floor but to take one small step in the steepest downward direction, then feel the ground again. Steps too large will overshoot the valley — possibly climbing higher; steps too small will never arrive before dark. The "learning rate" is simply how far you step.
The geometry of gradient descent: each step follows the steepest descent from where you stand, scaled by the learning rate
The learning rate decides success: three runs share identical data and initial weights, differing only in step size (click the legend to toggle a curve)
- Too large (oscillates, diverges)
- Appropriate (converges steadily)
- Too small (barely moves)
طريقة العمل
- 01
Forward: measure how wrong you currently are
Push a batch of samples through the model, obtain predictions, compare them with the truth, and reduce the difference to a single scalar loss. That scalar is the elevation of the hill you are standing on.
- 02
Backward: automatic differentiation gives the direction
Backpropagation walks the computation graph from output back to input, applying the chain rule layer by layer, obtaining in one pass the partial derivative of the loss with respect to every parameter. This is the central trick of modern frameworks: turning "derive the gradient" from mathematical drudgery into a single function call.
- 03
Update: take the step
Parameter ← parameter − learning rate × gradient. One step looks trivial, yet after millions of them a randomly initialised network grows internal structure that can recognise speech and write code.
- 04
Momentum and adaptivity: making the step smarter
Plain gradient descent zig-zags inside a narrow valley. Momentum accumulates past directions to smooth the oscillation, while Adam tunes the step size per parameter — the default optimiser for essentially every large model today.
الصيغة الأساسية
θ ← θ − η · ∇θ L(θ)The typical character of three optimisers: convergence speed, stability, sensitivity to hyperparameters and memory cost
- SGD + momentum
- Adam
- RMSProp
مجالات الاستخدام
- Training every neural network: from handwritten-digit recognition to hundred-billion-parameter language models
- Fine-tuning: continue descending on a new task with a tiny learning rate so existing abilities are not destroyed
- Reinforcement learning: treat policy performance as the loss, so gradient ascent becomes policy improvement
- Adversarial examples: perturbing the input along the gradient can flip a model’s judgement entirely
مفاهيم خاطئة شائعة
- A gradient is purely local. It tells you the steepest direction underfoot, not whether that direction leads to the global optimum. Deep losses teem with local valleys and saddle points, yet in practice sufficiently wide networks are full of "good enough" local solutions.
- The learning rate is the hyperparameter that matters most. On identical data, a tenfold difference can separate "converges beautifully" from "loss becomes NaN".
- Vanishing and exploding gradients: the chain rule multiplies terms, so gradients in deep networks shrink or blow up exponentially — which is precisely why residual connections, normalisation and gradient clipping exist.
مصطلحات أساسية
- Learning rate η
- How far each step moves
- Batch size
- How many samples estimate the gradient per step
- SGD
- Approximating the full gradient with a mini-batch
- Loss surface
- The high-dimensional terrain of loss values over parameter space