Hàm kích hoạt
Không có nó, mạng dù sâu đến đâu cũng chỉ là một phép biến đổi tuyến tính
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
ĐỊNH NGHĨA
An activation function is the nonlinear function applied after a neuron’s weighted sum; it decides how strongly that neuron fires. Common choices are Sigmoid, Tanh, ReLU and its variants. It is exactly what makes a multilayer network more than a single layer: remove the nonlinearity and any stack of linear maps collapses into one linear map, leaving depth without purpose.
Trực giác
Picture a row of switches with thresholds. The linear part sets a voltage; the activation decides whether that voltage actually lights the lamp — too low and it stays fully dark (ReLU is zero on the negative side), too high and it saturates (Sigmoid barely changes near its ends). This bend — either switch off completely or pass through with a fixed slope — is what lets many simple weighted sums assemble into arbitrarily complex shapes.
A shape comparison of four activations: the step function (original perceptron), Sigmoid and Tanh saturating at both ends, and ReLU staying linear on the positive side
- Step
- Sigmoid
- Tanh
- ReLU
Top-1 accuracy on CIFAR-10 for the same convolutional network and training setup, replacing only the activation: modern nonlinearities (ReLU/GELU) clearly beat the saturating ones (magnitudes from typical reported results)
Cách hoạt động
- 01
Sigmoid and the saturation problem
σ(z) = 1/(1+e⁻ᶻ) squashes any real number into (0,1). It is smooth and can emit probabilities directly, but its derivative vanishes at both ends, so stacking it deep makes gradients decay exponentially.
- 02
Tanh and zero-centring
tanh outputs in (−1,1) and is zero-centred, which trains better than Sigmoid; yet it still saturates, so deep networks remain hard to train.
- 03
The ReLU breakthrough
ReLU(z) = max(0, z) has derivative 1 on the positive side, so gradients pass cleanly through many layers — fast to train and cheap to compute. The price is a fully dead negative side, which can leave neurons permanently inactive ("dying ReLU").
- 04
Variants and smooth gating
LeakyReLU gives the negative side a small slope to revive dying neurons, while GELU uses a Gaussian CDF as a smooth gate. Being smooth and differentiable, GELU tends to perform better and has become the default in Transformers.
Ứng dụng
- Introducing nonlinearity in hidden layers, most often with ReLU or GELU
- Matching the output to the task: Sigmoid for binary, Softmax for multiclass
- Inside gating units: LSTM/GRU use Sigmoid as an on-off switch
- Smooth gating in generative models and attention (GELU/SiLU)
Hiểu lầm thường gặp
- A common misconception is that "depth itself brings nonlinearity". The opposite holds: without activation functions, a multilayer network is mathematically equivalent to a single linear model.
- Sigmoid’s output is not zero-centred, which makes the optimisation path zig-zag; this, besides saturation, is one reason it fell out of favour in deep networks.
- Dying ReLUs and the zero gradient on the negative side are real problems; LeakyReLU and GELU are designed to mitigate them rather than as mere performance embellishments.
Thuật ngữ chính
- Saturation
- A function whose derivative tends to 0 at the extremes, blocking gradients
- Dying ReLU
- A neuron stuck in the negative region with zero gradient, no longer updating
- Zero-centred
- Outputs symmetric about 0, which aids optimisation
- Vanishing gradient
- Gradients shrinking exponentially as they are multiplied across layers