Ir para o conteúdo
Atlas de IA

Normalização e conexões residuais

Tornar treinável uma rede de cem camadas: um atalho de identidade e uma reescala por camada

03 Aprendizado profundoEspecialistaEntrada 6 deste domínio

O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.

DEFINIÇÃO

Normalisation (BatchNorm/LayerNorm/RMSNorm) rescales activations at each layer to stabilise the distribution of inputs to subsequent layers; residual connections let each layer learn only the change on top of its input (H(x) = F(x) + x) and give gradients an identity shortcut. Together they make deep networks trainable that would otherwise degrade or fail to converge.

Intuição

A residual connection is like adding a direct line to a long relay: even if a few batons are dropped, the signal still arrives intact along the shortcut, and gradients need not traverse every layer. Normalisation is like re-standardising the raw material before each station starts — without it, the deeper the network the more the values drift, and training proceeds as if measuring with two mismatched rulers. The two address two faces of one problem: how a signal can flow through a very deep network without decaying or drifting.

Fig. 1

Plain stacking versus residual connections: the former forces gradients through every layer, the latter offers an identity shortcut that keeps hundred-layer networks trainable

Fig. 2

Two fates of going deeper on ImageNet: plain stacking degrades (deeper is worse), while normalisation plus residuals make deeper more accurate. Values are top-1 accuracy (%), from the order of magnitude reported in the ResNet paper

Como funciona

  1. 01

    Internal covariate shift

    Whenever one layer updates, the input distribution of every later layer shifts, forcing them to re-adapt. Normalisation re-centres and re-scales activations to a stable mean and variance, easing this "chasing a moving target" problem.

  2. 02

    BatchNorm / LayerNorm / RMSNorm

    BatchNorm normalises across the batch dimension — effective for CNNs but batch-dependent and ill-suited to variable-length sequences. LayerNorm normalises over the feature dimension of each sample, is batch-independent, and underpins Transformers. RMSNorm drops the mean and scales only by the root mean square, which is faster and cheaper.

  3. 03

    Residual connections and the degradation problem

    Deeper plain networks show higher training error — degradation, not overfitting. Residuals make the identity map trivial to represent (just drive F to zero), so going deeper can at least never hurt.

  4. 04

    He initialisation and gradient scale

    For ReLU, He initialisation sets the weight variance to 2/n (n being the fan-in), keeping the variance of forward activations and backward gradients stable across layers — the necessary companion that lets residual networks stack many blocks.

Onde é usado

  • Residuals plus normalisation are standard in ResNet and nearly every deep CNN
  • LayerNorm/RMSNorm are core components of Transformers and large language models
  • Stabilising large-scale distributed training: faster convergence and larger learning rates
  • Combined with He initialisation and pre-activation blocks to further improve stability

Equívocos comuns

  • BatchNorm’s statistics are unstable with small batches; at inference it must use running averages accumulated during training, otherwise results fluctuate wildly with batch size.
  • LayerNorm/RMSNorm and BatchNorm are not freely interchangeable: the former suit sequences and variable-length data, the latter are often more effective on convolutions and large-batch image tasks.
  • Residuals mitigate degradation but do not replace regularisation; very deep models still need careful tuning or they can remain unstable.

Termos-chave

Internal covariate shift
The shifting distribution of inputs to later layers during training
Degradation problem
Deeper networks with higher training error, and not from overfitting
Identity shortcut
The path in a residual connection that adds the input straight back to the output
Pre-activation
A layout placing normalisation before the convolution, which trains more stably

Leituras complementares