Aller au contenu
Atlas de l'IA

Neurone et perceptron

La plus petite pièce d’un réseau de neurones : une somme pondérée, un biais et une touche de non-linéarité

03 Apprentissage profondDébutantEntrée 1 de ce domaine

Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.

DÉFINITION

A neuron is the basic computational unit of a neural network: it takes several inputs, assigns each a weight, sums them, adds a bias, and passes the result through an activation function. The single-layer perceptron, the earliest such model, computes one weighted sum and thresholds it into 0 or 1, so geometrically it can only separate data with a single line (or hyperplane).

Intuition

Think of it as scoring. To decide whether an email is spam, give each clue a weight: "congratulations, you won" adds 0.8, a colleague’s signature subtracts 0.5, and so on. Sum the weighted clues, add a baseline bias, and call it spam once the total crosses a threshold. The perceptron’s limitation is just as concrete: it can only draw a single straight line, so a problem like XOR — where the two classes interleave — cannot be separated wherever the line is placed. Expressing such a relation requires inserting a new layer of neurons in between.

Fig. 1

The anatomy of a neuron: three inputs each carry a weight; they are summed, offset by a bias, and passed through an activation to produce the output

Fig. 2

A single-layer perceptron cannot learn XOR: its loss stays pinned near 0.69 — the chance level — while a multilayer network with a hidden layer drives it towards zero (click the legend to toggle)

  • Single-layer perceptron (cannot converge)
  • Multilayer network with a hidden layer

Fonctionnement

  1. 01

    Weighted sum: assign a weight to each clue

    Multiply each input xᵢ by its weight wᵢ and add everything up. A weight says how much a clue matters; it may be positive or negative, and it is what training learns.

  2. 02

    Add a bias: shift the threshold

    Add a bias b, a learnable offset applied to the decision threshold. Without it, a neuron produces a fixed value whenever all inputs are zero, which sharply reduces its flexibility.

  3. 03

    Through an activation: introduce nonlinearity

    The weighted sum z = Σ wᵢxᵢ + b is fed into an activation function — a step function (0 or 1) in the early perceptron, a smooth function such as ReLU today. This step is what gives the neuron its ability to switch or to bend.

  4. 04

    The single-layer limit: from perceptron to multilayer networks

    A single perceptron can only report which side of a line a point falls on. Expressing nonlinear relations such as XOR requires stacking neurons into a hidden layer that first maps the inputs into a new space — one in which the previously inseparable data becomes separable. This is precisely why multilayer networks exist.

Fig. 3

Why hidden layers are needed: a single perceptron has only one line and cannot separate XOR; a hidden layer maps the inputs into a new space where the same problem becomes linearly separable

Où c'est utilisé

  • The geometric basis of linear classification and logistic regression, and the starting point for reasoning about decision boundaries
  • The building block of deep networks: a fully connected layer is many such neurons side by side
  • Interpretable scorecards: weighted feature sums in credit scoring and simple recommenders
  • A minimal model for theory: discussing linear separability, capacity and convergence

Idées fausses courantes

  • A single-layer perceptron cannot learn XOR not because training is too short but because of a hard capacity limit — no number of iterations or placement of the line will ever separate the classes.
  • It is tempting to think that "more layers means nonlinearity", yet before activation functions are added, a stack of linear maps still collapses into a single linear map (see the entry on activation functions).
  • A large parameter count does not imply the ability to fit arbitrary functions. Universal approximation requires width, nonlinearity and sufficient depth together.

Termes clés

Weight
How strongly an input influences the output; may be positive or negative
Bias
A learnable offset applied to the threshold
Activation function
A function that applies a nonlinear transform to the weighted sum
Linearly separable
A hyperplane exists that separates the two classes perfectly

Lectures complémentaires