تخطٍّ إلى المحتوى
أطلس الذكاء الاصطناعي

دوال التنشيط

لولاها، لكانت أعمق شبكة مجرّد تحويل خطي واحد

03 التعلّم العميقمبتدئالمدخل 3 في هذا المجال

يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.

التعريف

An activation function is the nonlinear function applied after a neuron’s weighted sum; it decides how strongly that neuron fires. Common choices are Sigmoid, Tanh, ReLU and its variants. It is exactly what makes a multilayer network more than a single layer: remove the nonlinearity and any stack of linear maps collapses into one linear map, leaving depth without purpose.

الحدس المباشر

Picture a row of switches with thresholds. The linear part sets a voltage; the activation decides whether that voltage actually lights the lamp — too low and it stays fully dark (ReLU is zero on the negative side), too high and it saturates (Sigmoid barely changes near its ends). This bend — either switch off completely or pass through with a fixed slope — is what lets many simple weighted sums assemble into arbitrarily complex shapes.

شكل 1

A shape comparison of four activations: the step function (original perceptron), Sigmoid and Tanh saturating at both ends, and ReLU staying linear on the positive side

  • Step
  • Sigmoid
  • Tanh
  • ReLU
شكل 2

Top-1 accuracy on CIFAR-10 for the same convolutional network and training setup, replacing only the activation: modern nonlinearities (ReLU/GELU) clearly beat the saturating ones (magnitudes from typical reported results)

طريقة العمل

  1. 01

    Sigmoid and the saturation problem

    σ(z) = 1/(1+e⁻ᶻ) squashes any real number into (0,1). It is smooth and can emit probabilities directly, but its derivative vanishes at both ends, so stacking it deep makes gradients decay exponentially.

  2. 02

    Tanh and zero-centring

    tanh outputs in (−1,1) and is zero-centred, which trains better than Sigmoid; yet it still saturates, so deep networks remain hard to train.

  3. 03

    The ReLU breakthrough

    ReLU(z) = max(0, z) has derivative 1 on the positive side, so gradients pass cleanly through many layers — fast to train and cheap to compute. The price is a fully dead negative side, which can leave neurons permanently inactive ("dying ReLU").

  4. 04

    Variants and smooth gating

    LeakyReLU gives the negative side a small slope to revive dying neurons, while GELU uses a Gaussian CDF as a smooth gate. Being smooth and differentiable, GELU tends to perform better and has become the default in Transformers.

مجالات الاستخدام

  • Introducing nonlinearity in hidden layers, most often with ReLU or GELU
  • Matching the output to the task: Sigmoid for binary, Softmax for multiclass
  • Inside gating units: LSTM/GRU use Sigmoid as an on-off switch
  • Smooth gating in generative models and attention (GELU/SiLU)

مفاهيم خاطئة شائعة

  • A common misconception is that "depth itself brings nonlinearity". The opposite holds: without activation functions, a multilayer network is mathematically equivalent to a single linear model.
  • Sigmoid’s output is not zero-centred, which makes the optimisation path zig-zag; this, besides saturation, is one reason it fell out of favour in deep networks.
  • Dying ReLUs and the zero gradient on the negative side are real problems; LeakyReLU and GELU are designed to mitigate them rather than as mere performance embellishments.

مصطلحات أساسية

Saturation
A function whose derivative tends to 0 at the extremes, blocking gradients
Dying ReLU
A neuron stuck in the negative region with zero gradient, no longer updating
Zero-centred
Outputs symmetric about 0, which aids optimisation
Vanishing gradient
Gradients shrinking exponentially as they are multiplied across layers

قراءات موسّعة