Funciones de activación
Sin ella, hasta la red más profunda es solo una única transformación lineal
El texto completo se presenta en inglés; el título y el resumen están traducidos.
DEFINICIÓN
An activation function is the nonlinear function applied after a neuron’s weighted sum; it decides how strongly that neuron fires. Common choices are Sigmoid, Tanh, ReLU and its variants. It is exactly what makes a multilayer network more than a single layer: remove the nonlinearity and any stack of linear maps collapses into one linear map, leaving depth without purpose.
Intuición
Picture a row of switches with thresholds. The linear part sets a voltage; the activation decides whether that voltage actually lights the lamp — too low and it stays fully dark (ReLU is zero on the negative side), too high and it saturates (Sigmoid barely changes near its ends). This bend — either switch off completely or pass through with a fixed slope — is what lets many simple weighted sums assemble into arbitrarily complex shapes.
A shape comparison of four activations: the step function (original perceptron), Sigmoid and Tanh saturating at both ends, and ReLU staying linear on the positive side
- Step
- Sigmoid
- Tanh
- ReLU
Top-1 accuracy on CIFAR-10 for the same convolutional network and training setup, replacing only the activation: modern nonlinearities (ReLU/GELU) clearly beat the saturating ones (magnitudes from typical reported results)
Cómo funciona
- 01
Sigmoid and the saturation problem
σ(z) = 1/(1+e⁻ᶻ) squashes any real number into (0,1). It is smooth and can emit probabilities directly, but its derivative vanishes at both ends, so stacking it deep makes gradients decay exponentially.
- 02
Tanh and zero-centring
tanh outputs in (−1,1) and is zero-centred, which trains better than Sigmoid; yet it still saturates, so deep networks remain hard to train.
- 03
The ReLU breakthrough
ReLU(z) = max(0, z) has derivative 1 on the positive side, so gradients pass cleanly through many layers — fast to train and cheap to compute. The price is a fully dead negative side, which can leave neurons permanently inactive ("dying ReLU").
- 04
Variants and smooth gating
LeakyReLU gives the negative side a small slope to revive dying neurons, while GELU uses a Gaussian CDF as a smooth gate. Being smooth and differentiable, GELU tends to perform better and has become the default in Transformers.
Dónde se usa
- Introducing nonlinearity in hidden layers, most often with ReLU or GELU
- Matching the output to the task: Sigmoid for binary, Softmax for multiclass
- Inside gating units: LSTM/GRU use Sigmoid as an on-off switch
- Smooth gating in generative models and attention (GELU/SiLU)
Errores comunes
- A common misconception is that "depth itself brings nonlinearity". The opposite holds: without activation functions, a multilayer network is mathematically equivalent to a single linear model.
- Sigmoid’s output is not zero-centred, which makes the optimisation path zig-zag; this, besides saturation, is one reason it fell out of favour in deep networks.
- Dying ReLUs and the zero gradient on the negative side are real problems; LeakyReLU and GELU are designed to mitigate them rather than as mere performance embellishments.
Términos clave
- Saturation
- A function whose derivative tends to 0 at the extremes, blocking gradients
- Dying ReLU
- A neuron stuck in the negative region with zero gradient, no longer updating
- Zero-centred
- Outputs symmetric about 0, which aids optimisation
- Vanishing gradient
- Gradients shrinking exponentially as they are multiplied across layers