본문으로 건너뛰기
AI 도감

활성화 함수

이것이 없으면 아무리 깊은 네트워크도 단 한 번의 선형 변환일 뿐이다

03 딥러닝입문이 영역의 3번째 항목

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

정의

An activation function is the nonlinear function applied after a neuron’s weighted sum; it decides how strongly that neuron fires. Common choices are Sigmoid, Tanh, ReLU and its variants. It is exactly what makes a multilayer network more than a single layer: remove the nonlinearity and any stack of linear maps collapses into one linear map, leaving depth without purpose.

직관적 이해

Picture a row of switches with thresholds. The linear part sets a voltage; the activation decides whether that voltage actually lights the lamp — too low and it stays fully dark (ReLU is zero on the negative side), too high and it saturates (Sigmoid barely changes near its ends). This bend — either switch off completely or pass through with a fixed slope — is what lets many simple weighted sums assemble into arbitrarily complex shapes.

그림 1

A shape comparison of four activations: the step function (original perceptron), Sigmoid and Tanh saturating at both ends, and ReLU staying linear on the positive side

  • Step
  • Sigmoid
  • Tanh
  • ReLU
그림 2

Top-1 accuracy on CIFAR-10 for the same convolutional network and training setup, replacing only the activation: modern nonlinearities (ReLU/GELU) clearly beat the saturating ones (magnitudes from typical reported results)

작동 원리

  1. 01

    Sigmoid and the saturation problem

    σ(z) = 1/(1+e⁻ᶻ) squashes any real number into (0,1). It is smooth and can emit probabilities directly, but its derivative vanishes at both ends, so stacking it deep makes gradients decay exponentially.

  2. 02

    Tanh and zero-centring

    tanh outputs in (−1,1) and is zero-centred, which trains better than Sigmoid; yet it still saturates, so deep networks remain hard to train.

  3. 03

    The ReLU breakthrough

    ReLU(z) = max(0, z) has derivative 1 on the positive side, so gradients pass cleanly through many layers — fast to train and cheap to compute. The price is a fully dead negative side, which can leave neurons permanently inactive ("dying ReLU").

  4. 04

    Variants and smooth gating

    LeakyReLU gives the negative side a small slope to revive dying neurons, while GELU uses a Gaussian CDF as a smooth gate. Being smooth and differentiable, GELU tends to perform better and has become the default in Transformers.

응용 분야

  • Introducing nonlinearity in hidden layers, most often with ReLU or GELU
  • Matching the output to the task: Sigmoid for binary, Softmax for multiclass
  • Inside gating units: LSTM/GRU use Sigmoid as an on-off switch
  • Smooth gating in generative models and attention (GELU/SiLU)

흔한 오해

  • A common misconception is that "depth itself brings nonlinearity". The opposite holds: without activation functions, a multilayer network is mathematically equivalent to a single linear model.
  • Sigmoid’s output is not zero-centred, which makes the optimisation path zig-zag; this, besides saturation, is one reason it fell out of favour in deep networks.
  • Dying ReLUs and the zero gradient on the negative side are real problems; LeakyReLU and GELU are designed to mitigate them rather than as mere performance embellishments.

핵심 용어

Saturation
A function whose derivative tends to 0 at the extremes, blocking gradients
Dying ReLU
A neuron stuck in the negative region with zero gradient, no longer updating
Zero-centred
Outputs symmetric about 0, which aids optimisation
Vanishing gradient
Gradients shrinking exponentially as they are multiplied across layers

참고문헌