Chuyển đến nội dung
Bản đồ AI

Chuẩn hóa và kết nối dư

Làm mạng trăm tầng thực sự huấn luyện được: một lối tắt đồng nhất và một lần hiệu chỉnh lại ở mỗi tầng

03 Học sâuChuyên giaMục từ thứ 6 trong lĩnh vực

Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.

ĐỊNH NGHĨA

Normalisation (BatchNorm/LayerNorm/RMSNorm) rescales activations at each layer to stabilise the distribution of inputs to subsequent layers; residual connections let each layer learn only the change on top of its input (H(x) = F(x) + x) and give gradients an identity shortcut. Together they make deep networks trainable that would otherwise degrade or fail to converge.

Trực giác

A residual connection is like adding a direct line to a long relay: even if a few batons are dropped, the signal still arrives intact along the shortcut, and gradients need not traverse every layer. Normalisation is like re-standardising the raw material before each station starts — without it, the deeper the network the more the values drift, and training proceeds as if measuring with two mismatched rulers. The two address two faces of one problem: how a signal can flow through a very deep network without decaying or drifting.

Hình 1

Plain stacking versus residual connections: the former forces gradients through every layer, the latter offers an identity shortcut that keeps hundred-layer networks trainable

Hình 2

Two fates of going deeper on ImageNet: plain stacking degrades (deeper is worse), while normalisation plus residuals make deeper more accurate. Values are top-1 accuracy (%), from the order of magnitude reported in the ResNet paper

Cách hoạt động

  1. 01

    Internal covariate shift

    Whenever one layer updates, the input distribution of every later layer shifts, forcing them to re-adapt. Normalisation re-centres and re-scales activations to a stable mean and variance, easing this "chasing a moving target" problem.

  2. 02

    BatchNorm / LayerNorm / RMSNorm

    BatchNorm normalises across the batch dimension — effective for CNNs but batch-dependent and ill-suited to variable-length sequences. LayerNorm normalises over the feature dimension of each sample, is batch-independent, and underpins Transformers. RMSNorm drops the mean and scales only by the root mean square, which is faster and cheaper.

  3. 03

    Residual connections and the degradation problem

    Deeper plain networks show higher training error — degradation, not overfitting. Residuals make the identity map trivial to represent (just drive F to zero), so going deeper can at least never hurt.

  4. 04

    He initialisation and gradient scale

    For ReLU, He initialisation sets the weight variance to 2/n (n being the fan-in), keeping the variance of forward activations and backward gradients stable across layers — the necessary companion that lets residual networks stack many blocks.

Ứng dụng

  • Residuals plus normalisation are standard in ResNet and nearly every deep CNN
  • LayerNorm/RMSNorm are core components of Transformers and large language models
  • Stabilising large-scale distributed training: faster convergence and larger learning rates
  • Combined with He initialisation and pre-activation blocks to further improve stability

Hiểu lầm thường gặp

  • BatchNorm’s statistics are unstable with small batches; at inference it must use running averages accumulated during training, otherwise results fluctuate wildly with batch size.
  • LayerNorm/RMSNorm and BatchNorm are not freely interchangeable: the former suit sequences and variable-length data, the latter are often more effective on convolutions and large-batch image tasks.
  • Residuals mitigate degradation but do not replace regularisation; very deep models still need careful tuning or they can remain unstable.

Thuật ngữ chính

Internal covariate shift
The shifting distribution of inputs to later layers during training
Degradation problem
Deeper networks with higher training error, and not from overfitting
Identity shortcut
The path in a residual connection that adds the input straight back to the output
Pre-activation
A layout placing normalisation before the convolution, which trains more stably

Đọc thêm