Zum Inhalt springen
KI-Atlas

Normalisierung und Residualverbindungen

Ein Netz mit hundert Schichten wirklich trainierbar machen: ein Identitäts-Kurzweg plus eine Neuskalierung pro Schicht

03 Deep LearningExperte6. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

Normalisation (BatchNorm/LayerNorm/RMSNorm) rescales activations at each layer to stabilise the distribution of inputs to subsequent layers; residual connections let each layer learn only the change on top of its input (H(x) = F(x) + x) and give gradients an identity shortcut. Together they make deep networks trainable that would otherwise degrade or fail to converge.

Intuition

A residual connection is like adding a direct line to a long relay: even if a few batons are dropped, the signal still arrives intact along the shortcut, and gradients need not traverse every layer. Normalisation is like re-standardising the raw material before each station starts — without it, the deeper the network the more the values drift, and training proceeds as if measuring with two mismatched rulers. The two address two faces of one problem: how a signal can flow through a very deep network without decaying or drifting.

Abb. 1

Plain stacking versus residual connections: the former forces gradients through every layer, the latter offers an identity shortcut that keeps hundred-layer networks trainable

Abb. 2

Two fates of going deeper on ImageNet: plain stacking degrades (deeper is worse), while normalisation plus residuals make deeper more accurate. Values are top-1 accuracy (%), from the order of magnitude reported in the ResNet paper

Funktionsweise

  1. 01

    Internal covariate shift

    Whenever one layer updates, the input distribution of every later layer shifts, forcing them to re-adapt. Normalisation re-centres and re-scales activations to a stable mean and variance, easing this "chasing a moving target" problem.

  2. 02

    BatchNorm / LayerNorm / RMSNorm

    BatchNorm normalises across the batch dimension — effective for CNNs but batch-dependent and ill-suited to variable-length sequences. LayerNorm normalises over the feature dimension of each sample, is batch-independent, and underpins Transformers. RMSNorm drops the mean and scales only by the root mean square, which is faster and cheaper.

  3. 03

    Residual connections and the degradation problem

    Deeper plain networks show higher training error — degradation, not overfitting. Residuals make the identity map trivial to represent (just drive F to zero), so going deeper can at least never hurt.

  4. 04

    He initialisation and gradient scale

    For ReLU, He initialisation sets the weight variance to 2/n (n being the fan-in), keeping the variance of forward activations and backward gradients stable across layers — the necessary companion that lets residual networks stack many blocks.

Anwendungsfelder

  • Residuals plus normalisation are standard in ResNet and nearly every deep CNN
  • LayerNorm/RMSNorm are core components of Transformers and large language models
  • Stabilising large-scale distributed training: faster convergence and larger learning rates
  • Combined with He initialisation and pre-activation blocks to further improve stability

Häufige Missverständnisse

  • BatchNorm’s statistics are unstable with small batches; at inference it must use running averages accumulated during training, otherwise results fluctuate wildly with batch size.
  • LayerNorm/RMSNorm and BatchNorm are not freely interchangeable: the former suit sequences and variable-length data, the latter are often more effective on convolutions and large-batch image tasks.
  • Residuals mitigate degradation but do not replace regularisation; very deep models still need careful tuning or they can remain unstable.

Schlüsselbegriffe

Internal covariate shift
The shifting distribution of inputs to later layers during training
Degradation problem
Deeper networks with higher training error, and not from overfitting
Identity shortcut
The path in a residual connection that adds the input straight back to the output
Pre-activation
A layout placing normalisation before the convolution, which trains more stably

Weiterführende Literatur