सामान्यीकरण और अवशिष्ट संयोजन
सैकड़ों परतों वाले जाल को वाकई प्रशिक्षित करना: एक तत्समक शॉर्टकट और हर परत पर पुनः-मापन
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
परिभाषा
Normalisation (BatchNorm/LayerNorm/RMSNorm) rescales activations at each layer to stabilise the distribution of inputs to subsequent layers; residual connections let each layer learn only the change on top of its input (H(x) = F(x) + x) and give gradients an identity shortcut. Together they make deep networks trainable that would otherwise degrade or fail to converge.
सहज समझ
A residual connection is like adding a direct line to a long relay: even if a few batons are dropped, the signal still arrives intact along the shortcut, and gradients need not traverse every layer. Normalisation is like re-standardising the raw material before each station starts — without it, the deeper the network the more the values drift, and training proceeds as if measuring with two mismatched rulers. The two address two faces of one problem: how a signal can flow through a very deep network without decaying or drifting.
Plain stacking versus residual connections: the former forces gradients through every layer, the latter offers an identity shortcut that keeps hundred-layer networks trainable
Two fates of going deeper on ImageNet: plain stacking degrades (deeper is worse), while normalisation plus residuals make deeper more accurate. Values are top-1 accuracy (%), from the order of magnitude reported in the ResNet paper
कार्यप्रणाली
- 01
Internal covariate shift
Whenever one layer updates, the input distribution of every later layer shifts, forcing them to re-adapt. Normalisation re-centres and re-scales activations to a stable mean and variance, easing this "chasing a moving target" problem.
- 02
BatchNorm / LayerNorm / RMSNorm
BatchNorm normalises across the batch dimension — effective for CNNs but batch-dependent and ill-suited to variable-length sequences. LayerNorm normalises over the feature dimension of each sample, is batch-independent, and underpins Transformers. RMSNorm drops the mean and scales only by the root mean square, which is faster and cheaper.
- 03
Residual connections and the degradation problem
Deeper plain networks show higher training error — degradation, not overfitting. Residuals make the identity map trivial to represent (just drive F to zero), so going deeper can at least never hurt.
- 04
He initialisation and gradient scale
For ReLU, He initialisation sets the weight variance to 2/n (n being the fan-in), keeping the variance of forward activations and backward gradients stable across layers — the necessary companion that lets residual networks stack many blocks.
उपयोग के क्षेत्र
- Residuals plus normalisation are standard in ResNet and nearly every deep CNN
- LayerNorm/RMSNorm are core components of Transformers and large language models
- Stabilising large-scale distributed training: faster convergence and larger learning rates
- Combined with He initialisation and pre-activation blocks to further improve stability
सामान्य भ्रांतियाँ
- BatchNorm’s statistics are unstable with small batches; at inference it must use running averages accumulated during training, otherwise results fluctuate wildly with batch size.
- LayerNorm/RMSNorm and BatchNorm are not freely interchangeable: the former suit sequences and variable-length data, the latter are often more effective on convolutions and large-batch image tasks.
- Residuals mitigate degradation but do not replace regularisation; very deep models still need careful tuning or they can remain unstable.
मुख्य शब्द
- Internal covariate shift
- The shifting distribution of inputs to later layers during training
- Degradation problem
- Deeper networks with higher training error, and not from overfitting
- Identity shortcut
- The path in a residual connection that adds the input straight back to the output
- Pre-activation
- A layout placing normalisation before the convolution, which trains more stably