Zum Inhalt springen
KI-Atlas

Aktivierungsfunktionen

Ohne sie ist selbst das tiefste Netz nur eine einzige lineare Abbildung

03 Deep LearningAnfänger3. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

An activation function is the nonlinear function applied after a neuron’s weighted sum; it decides how strongly that neuron fires. Common choices are Sigmoid, Tanh, ReLU and its variants. It is exactly what makes a multilayer network more than a single layer: remove the nonlinearity and any stack of linear maps collapses into one linear map, leaving depth without purpose.

Intuition

Picture a row of switches with thresholds. The linear part sets a voltage; the activation decides whether that voltage actually lights the lamp — too low and it stays fully dark (ReLU is zero on the negative side), too high and it saturates (Sigmoid barely changes near its ends). This bend — either switch off completely or pass through with a fixed slope — is what lets many simple weighted sums assemble into arbitrarily complex shapes.

Abb. 1

A shape comparison of four activations: the step function (original perceptron), Sigmoid and Tanh saturating at both ends, and ReLU staying linear on the positive side

  • Step
  • Sigmoid
  • Tanh
  • ReLU
Abb. 2

Top-1 accuracy on CIFAR-10 for the same convolutional network and training setup, replacing only the activation: modern nonlinearities (ReLU/GELU) clearly beat the saturating ones (magnitudes from typical reported results)

Funktionsweise

  1. 01

    Sigmoid and the saturation problem

    σ(z) = 1/(1+e⁻ᶻ) squashes any real number into (0,1). It is smooth and can emit probabilities directly, but its derivative vanishes at both ends, so stacking it deep makes gradients decay exponentially.

  2. 02

    Tanh and zero-centring

    tanh outputs in (−1,1) and is zero-centred, which trains better than Sigmoid; yet it still saturates, so deep networks remain hard to train.

  3. 03

    The ReLU breakthrough

    ReLU(z) = max(0, z) has derivative 1 on the positive side, so gradients pass cleanly through many layers — fast to train and cheap to compute. The price is a fully dead negative side, which can leave neurons permanently inactive ("dying ReLU").

  4. 04

    Variants and smooth gating

    LeakyReLU gives the negative side a small slope to revive dying neurons, while GELU uses a Gaussian CDF as a smooth gate. Being smooth and differentiable, GELU tends to perform better and has become the default in Transformers.

Anwendungsfelder

  • Introducing nonlinearity in hidden layers, most often with ReLU or GELU
  • Matching the output to the task: Sigmoid for binary, Softmax for multiclass
  • Inside gating units: LSTM/GRU use Sigmoid as an on-off switch
  • Smooth gating in generative models and attention (GELU/SiLU)

Häufige Missverständnisse

  • A common misconception is that "depth itself brings nonlinearity". The opposite holds: without activation functions, a multilayer network is mathematically equivalent to a single linear model.
  • Sigmoid’s output is not zero-centred, which makes the optimisation path zig-zag; this, besides saturation, is one reason it fell out of favour in deep networks.
  • Dying ReLUs and the zero gradient on the negative side are real problems; LeakyReLU and GELU are designed to mitigate them rather than as mere performance embellishments.

Schlüsselbegriffe

Saturation
A function whose derivative tends to 0 at the extremes, blocking gradients
Dying ReLU
A neuron stuck in the negative region with zero gradient, no longer updating
Zero-centred
Outputs symmetric about 0, which aids optimisation
Vanishing gradient
Gradients shrinking exponentially as they are multiplied across layers

Weiterführende Literatur