Zum Inhalt springen
KI-Atlas

Neuron und Perzeptron

Das kleinste Teil eines neuronalen Netzes: eine gewichtete Summe, ein Bias und ein Hauch Nichtlinearität

03 Deep LearningAnfänger1. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

A neuron is the basic computational unit of a neural network: it takes several inputs, assigns each a weight, sums them, adds a bias, and passes the result through an activation function. The single-layer perceptron, the earliest such model, computes one weighted sum and thresholds it into 0 or 1, so geometrically it can only separate data with a single line (or hyperplane).

Intuition

Think of it as scoring. To decide whether an email is spam, give each clue a weight: "congratulations, you won" adds 0.8, a colleague’s signature subtracts 0.5, and so on. Sum the weighted clues, add a baseline bias, and call it spam once the total crosses a threshold. The perceptron’s limitation is just as concrete: it can only draw a single straight line, so a problem like XOR — where the two classes interleave — cannot be separated wherever the line is placed. Expressing such a relation requires inserting a new layer of neurons in between.

Abb. 1

The anatomy of a neuron: three inputs each carry a weight; they are summed, offset by a bias, and passed through an activation to produce the output

Abb. 2

A single-layer perceptron cannot learn XOR: its loss stays pinned near 0.69 — the chance level — while a multilayer network with a hidden layer drives it towards zero (click the legend to toggle)

  • Single-layer perceptron (cannot converge)
  • Multilayer network with a hidden layer

Funktionsweise

  1. 01

    Weighted sum: assign a weight to each clue

    Multiply each input xᵢ by its weight wᵢ and add everything up. A weight says how much a clue matters; it may be positive or negative, and it is what training learns.

  2. 02

    Add a bias: shift the threshold

    Add a bias b, a learnable offset applied to the decision threshold. Without it, a neuron produces a fixed value whenever all inputs are zero, which sharply reduces its flexibility.

  3. 03

    Through an activation: introduce nonlinearity

    The weighted sum z = Σ wᵢxᵢ + b is fed into an activation function — a step function (0 or 1) in the early perceptron, a smooth function such as ReLU today. This step is what gives the neuron its ability to switch or to bend.

  4. 04

    The single-layer limit: from perceptron to multilayer networks

    A single perceptron can only report which side of a line a point falls on. Expressing nonlinear relations such as XOR requires stacking neurons into a hidden layer that first maps the inputs into a new space — one in which the previously inseparable data becomes separable. This is precisely why multilayer networks exist.

Abb. 3

Why hidden layers are needed: a single perceptron has only one line and cannot separate XOR; a hidden layer maps the inputs into a new space where the same problem becomes linearly separable

Anwendungsfelder

  • The geometric basis of linear classification and logistic regression, and the starting point for reasoning about decision boundaries
  • The building block of deep networks: a fully connected layer is many such neurons side by side
  • Interpretable scorecards: weighted feature sums in credit scoring and simple recommenders
  • A minimal model for theory: discussing linear separability, capacity and convergence

Häufige Missverständnisse

  • A single-layer perceptron cannot learn XOR not because training is too short but because of a hard capacity limit — no number of iterations or placement of the line will ever separate the classes.
  • It is tempting to think that "more layers means nonlinearity", yet before activation functions are added, a stack of linear maps still collapses into a single linear map (see the entry on activation functions).
  • A large parameter count does not imply the ability to fit arbitrary functions. Universal approximation requires width, nonlinearity and sufficient depth together.

Schlüsselbegriffe

Weight
How strongly an input influences the output; may be positive or negative
Bias
A learnable offset applied to the threshold
Activation function
A function that applies a nonlinear transform to the weighted sum
Linearly separable
A hyperplane exists that separates the two classes perfectly

Weiterführende Literatur