الانتشار العكسي
تحويل «حساب التدرّج» من مشقّة رياضية إلى استدعاء دالة واحد — اللحظة التي انطلق فيها التعلّم العميق
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Backpropagation is an efficient algorithm for computing gradients. It first runs a forward pass over the computation graph to obtain the loss, then walks backward from the output to the inputs, applying the chain rule to derive the partial derivative of the loss with respect to every layer’s parameters. With it, a single backward sweep yields all gradients, no matter how many layers or parameters the network has.
الحدس المباشر
Picture a factory line: a product (the input) passes through stations (the layers) into a finished good (the prediction), and inspection measures the gap to spec (the loss). To trace how much each station should change, you walk back from the last station, distributing blame in proportion to contribution. Each station only needs the error returning from downstream, multiplies it by its own local derivative, and passes it further up — with no need to recompute the whole line. Backpropagation does exactly this, applying the chain rule precisely at every step.
The two sweeps of one training iteration: the solid path is forward (computing the loss), the dashed path is backward (propagating gradients by the chain rule)
Gradients decaying with depth: in an 8-layer network, Sigmoid makes layers near the input receive exponentially smaller gradients, whereas ReLU with normalised activations keeps them roughly even (log scale on the y-axis)
- Sigmoid (deep decay)
- ReLU + normalisation (stable)
طريقة العمل
- 01
Forward pass: obtain the loss and cache intermediates
Run the layers in order — a linear map followed by a nonlinearity — to obtain the prediction and a scalar loss, caching each layer’s intermediate activations for the backward pass.
- 02
Backward pass: the chain rule, layer by layer
Starting from the loss, multiply the gradient with respect to a layer’s output by that layer’s local derivative to obtain the gradient with respect to its input, then hand it to the previous layer. Each layer performs just one local multiplication.
- 03
Obtain parameter gradients: hand them to the optimiser
An outer product of the cached input with the incoming gradient (a convolution in convolutional networks) yields the gradients with respect to weights and biases, which a gradient-based optimiser consumes directly.
- 04
The meaning of autodiff: no more structural limits
A framework records any computation as a graph and performs backpropagation automatically. This is what freed researchers from hand-deriving partial derivatives and let them design arbitrarily deep, arbitrarily structured networks.
الصيغة الأساسية
∂L/∂W⁽ˡ⁾ = δ⁽ˡ⁾ (a⁽ˡ⁻¹⁾)ᵀ, δ⁽ˡ⁾ = ((W⁽ˡ⁺¹⁾)ᵀ δ⁽ˡ⁺¹⁾) ⊙ f′(z⁽ˡ⁾)مجالات الاستخدام
- The shared engine that trains every neural network — CNNs, RNNs and Transformers alike
- Bound to autodiff libraries (PyTorch, TensorFlow, JAX) to form the standard training loop
- Enables fine-tuning and parameter-efficient methods such as LoRA that keep backpropagating through existing weights
- Analysing gradient flow: diagnosing vanishing/exploding gradients or drawing saliency maps to explain predictions
مفاهيم خاطئة شائعة
- Backpropagation only computes gradients; it is distinct from the optimiser (SGD, Adam). The former yields a direction, the latter decides how to use it.
- Backpropagation must retain the forward activations, so memory grows with depth and batch size; gradient checkpointing trades recomputation for less storage.
- Gradients are local information. They give the steepest direction underfoot without any guarantee that it leads to the global optimum.
مصطلحات أساسية
- Computation graph
- A computation expressed as nodes and directed edges over which derivatives propagate
- Chain rule
- The derivative of a composition is the product of the local derivatives
- Automatic differentiation
- Letting a framework compute exact gradients automatically, not by numerical approximation
- Gradient flow
- The magnitude and stability of gradients as they propagate layer by layer