Neurone et perceptron
La plus petite pièce d’un réseau de neurones : une somme pondérée, un biais et une touche de non-linéarité
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
DÉFINITION
A neuron is the basic computational unit of a neural network: it takes several inputs, assigns each a weight, sums them, adds a bias, and passes the result through an activation function. The single-layer perceptron, the earliest such model, computes one weighted sum and thresholds it into 0 or 1, so geometrically it can only separate data with a single line (or hyperplane).
Intuition
Think of it as scoring. To decide whether an email is spam, give each clue a weight: "congratulations, you won" adds 0.8, a colleague’s signature subtracts 0.5, and so on. Sum the weighted clues, add a baseline bias, and call it spam once the total crosses a threshold. The perceptron’s limitation is just as concrete: it can only draw a single straight line, so a problem like XOR — where the two classes interleave — cannot be separated wherever the line is placed. Expressing such a relation requires inserting a new layer of neurons in between.
The anatomy of a neuron: three inputs each carry a weight; they are summed, offset by a bias, and passed through an activation to produce the output
A single-layer perceptron cannot learn XOR: its loss stays pinned near 0.69 — the chance level — while a multilayer network with a hidden layer drives it towards zero (click the legend to toggle)
- Single-layer perceptron (cannot converge)
- Multilayer network with a hidden layer
Fonctionnement
- 01
Weighted sum: assign a weight to each clue
Multiply each input xᵢ by its weight wᵢ and add everything up. A weight says how much a clue matters; it may be positive or negative, and it is what training learns.
- 02
Add a bias: shift the threshold
Add a bias b, a learnable offset applied to the decision threshold. Without it, a neuron produces a fixed value whenever all inputs are zero, which sharply reduces its flexibility.
- 03
Through an activation: introduce nonlinearity
The weighted sum z = Σ wᵢxᵢ + b is fed into an activation function — a step function (0 or 1) in the early perceptron, a smooth function such as ReLU today. This step is what gives the neuron its ability to switch or to bend.
- 04
The single-layer limit: from perceptron to multilayer networks
A single perceptron can only report which side of a line a point falls on. Expressing nonlinear relations such as XOR requires stacking neurons into a hidden layer that first maps the inputs into a new space — one in which the previously inseparable data becomes separable. This is precisely why multilayer networks exist.
Why hidden layers are needed: a single perceptron has only one line and cannot separate XOR; a hidden layer maps the inputs into a new space where the same problem becomes linearly separable
Où c'est utilisé
- The geometric basis of linear classification and logistic regression, and the starting point for reasoning about decision boundaries
- The building block of deep networks: a fully connected layer is many such neurons side by side
- Interpretable scorecards: weighted feature sums in credit scoring and simple recommenders
- A minimal model for theory: discussing linear separability, capacity and convergence
Idées fausses courantes
- A single-layer perceptron cannot learn XOR not because training is too short but because of a hard capacity limit — no number of iterations or placement of the line will ever separate the classes.
- It is tempting to think that "more layers means nonlinearity", yet before activation functions are added, a stack of linear maps still collapses into a single linear map (see the entry on activation functions).
- A large parameter count does not imply the ability to fit arbitrary functions. Universal approximation requires width, nonlinearity and sufficient depth together.
Termes clés
- Weight
- How strongly an input influences the output; may be positive or negative
- Bias
- A learnable offset applied to the threshold
- Activation function
- A function that applies a nonlinear transform to the weighted sum
- Linearly separable
- A hyperplane exists that separates the two classes perfectly