Zum Inhalt springen
KI-Atlas

Entropie und Informationstheorie

Je überraschender ein Satz, desto mehr Information trägt er – und die Verlustfunktion eines Sprachmodells misst genau diese Überraschung

01 Grundlagen der Mathematik und StatistikExperte6. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

Information theory measures information through uncertainty: the less likely an event, the more information it carries when it happens, and Shannon entropy is the average information of a random variable. Cross-entropy measures how much information it takes to encode data from one distribution using another, KL divergence is the excess, and perplexity is the language-model equivalent of cross-entropy.

Intuition

"The sun will rise tomorrow" carries almost no information because you expected it; "there will be a total solar eclipse tomorrow" carries a great deal. How much information a message carries depends on how much it surprises you. Training a language model is the same goal: make it as unsurprised as possible by the next piece of real text. The loss it reports is the average of exactly that surprise.

Abb. 1

Shannon’s model of communication: information content sets the limit of compression, and entropy is that limit

Abb. 2

Bernoulli entropy and cross-entropy: when the model encodes with a uniform q=0.5 the cross-entropy stays at 1 bit, and its gap above the true entropy H(p) is exactly the KL divergence (click the legend to toggle)

  • True entropy H(p)
  • Cross-entropy H(p, q=0.5)

Funktionsweise

  1. 01

    Information content: the more surprising, the more informative

    Define the information of an event as −log p: the smaller the probability p, the larger the information; the information of two independent events adds up, mirroring how their probabilities multiply. The logarithm is exactly what makes "information adds" match "probability multiplies", and with base 2 the unit is the bit.

  2. 02

    Entropy: an average measure of uncertainty

    Shannon entropy H(X) = −Σ p log p is the weighted average of the information of all values, measuring how hard the random variable is to guess. A fair coin has 1 bit of entropy, the hardest case; nearly certain events approach 0. Over a fixed value set, the more uniform the distribution, the larger the entropy.

  3. 03

    Cross-entropy: the cost of encoding reality with the wrong model

    Cross-entropy H(P,Q) = −Σ p(x) log q(x). Here the "true distribution" is P (the data) and the model supplies Q. It equals the true entropy H(P) plus the extra cost of using the wrong distribution. Training a language model is minimising cross-entropy — that is, minimising the model’s surprise at real text.

  4. 04

    KL divergence and perplexity: spelling out the cost and the hesitation

    KL divergence D(P‖Q) = H(P,Q) − H(P) is the part of the cross-entropy beyond the true entropy; it is non-negative and asymmetric, so it is not a distance. Perplexity = 2^{H(P,Q)} (or e^{H} with natural logs) reads as "how many equally likely options the model is hesitating among per step". In language-model papers, perplexity is often the more intuitive number.

Kernformel

H(P, Q) = − Σₓ P(x) log Q(x)
P is the true distribution (the data) and Q the model’s. The smaller the cross-entropy, the closer Q is to P; training minimises it.
Abb. 4

Perplexity of language models on WikiText-103: lower means the model is less surprised by real text (indicative magnitudes from public reports)

Anwendungsfelder

  • The training objective of language models: cross-entropy loss is negative log-likelihood
  • Evaluation metric: perplexity compares how well different models model text
  • Data compression: entropy is the theoretical limit on compression and guides the design of entropy coders
  • Decision trees split on "information gain" (a drop in entropy); variational inference minimises KL to approximate a posterior

Häufige Missverständnisse

  • Cross-entropy punishes confident mistakes very steeply: the more probability the model stakes on a wrong answer, the larger the loss. Conversely, underrating the correct answer is also costly — which is exactly what forces the model to be honest.
  • KL divergence is asymmetric, D(P‖Q) ≠ D(Q‖P), so it cannot be used as a distance; optimising the two directions gives very different results (one tends to cover, the other to lock onto modes).
  • Perplexity is comparable only under the same vocabulary and tokenisation. Comparing perplexity across models or tokenisers is badly misleading — finer tokenisation tends to lower perplexity without meaning a better model.

Schlüsselbegriffe

Self-information
The information of a single event, −log p
Shannon entropy
The average information, or uncertainty, of a random variable
Cross-entropy
The information needed to encode data from P using distribution Q
KL divergence
Cross-entropy minus true entropy; non-negative and asymmetric
Perplexity
The exponential of the cross-entropy; the effective number of options the model hesitates among per step

Weiterführende Literatur