본문으로 건너뛰기
AI 도감

정규화와 잔차 연결

수백 층 네트워크를 실제로 학습 가능하게: 항등 지름길 하나와 층마다의 재조정

03 딥러닝고급이 영역의 6번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

Normalisation (BatchNorm/LayerNorm/RMSNorm) rescales activations at each layer to stabilise the distribution of inputs to subsequent layers; residual connections let each layer learn only the change on top of its input (H(x) = F(x) + x) and give gradients an identity shortcut. Together they make deep networks trainable that would otherwise degrade or fail to converge.

직관적 이해

A residual connection is like adding a direct line to a long relay: even if a few batons are dropped, the signal still arrives intact along the shortcut, and gradients need not traverse every layer. Normalisation is like re-standardising the raw material before each station starts — without it, the deeper the network the more the values drift, and training proceeds as if measuring with two mismatched rulers. The two address two faces of one problem: how a signal can flow through a very deep network without decaying or drifting.

그림 1

Plain stacking versus residual connections: the former forces gradients through every layer, the latter offers an identity shortcut that keeps hundred-layer networks trainable

그림 2

Two fates of going deeper on ImageNet: plain stacking degrades (deeper is worse), while normalisation plus residuals make deeper more accurate. Values are top-1 accuracy (%), from the order of magnitude reported in the ResNet paper

작동 원리

  1. 01

    Internal covariate shift

    Whenever one layer updates, the input distribution of every later layer shifts, forcing them to re-adapt. Normalisation re-centres and re-scales activations to a stable mean and variance, easing this "chasing a moving target" problem.

  2. 02

    BatchNorm / LayerNorm / RMSNorm

    BatchNorm normalises across the batch dimension — effective for CNNs but batch-dependent and ill-suited to variable-length sequences. LayerNorm normalises over the feature dimension of each sample, is batch-independent, and underpins Transformers. RMSNorm drops the mean and scales only by the root mean square, which is faster and cheaper.

  3. 03

    Residual connections and the degradation problem

    Deeper plain networks show higher training error — degradation, not overfitting. Residuals make the identity map trivial to represent (just drive F to zero), so going deeper can at least never hurt.

  4. 04

    He initialisation and gradient scale

    For ReLU, He initialisation sets the weight variance to 2/n (n being the fan-in), keeping the variance of forward activations and backward gradients stable across layers — the necessary companion that lets residual networks stack many blocks.

응용 분야

  • Residuals plus normalisation are standard in ResNet and nearly every deep CNN
  • LayerNorm/RMSNorm are core components of Transformers and large language models
  • Stabilising large-scale distributed training: faster convergence and larger learning rates
  • Combined with He initialisation and pre-activation blocks to further improve stability

흔한 오해

  • BatchNorm’s statistics are unstable with small batches; at inference it must use running averages accumulated during training, otherwise results fluctuate wildly with batch size.
  • LayerNorm/RMSNorm and BatchNorm are not freely interchangeable: the former suit sequences and variable-length data, the latter are often more effective on convolutions and large-batch image tasks.
  • Residuals mitigate degradation but do not replace regularisation; very deep models still need careful tuning or they can remain unstable.

핵심 용어

Internal covariate shift
The shifting distribution of inputs to later layers during training
Degradation problem
Deeper networks with higher training error, and not from overfitting
Identity shortcut
The path in a residual connection that adds the input straight back to the output
Pre-activation
A layout placing normalisation before the convolution, which trains more stably

참고문헌