Entropie und Informationstheorie
Je überraschender ein Satz, desto mehr Information trägt er – und die Verlustfunktion eines Sprachmodells misst genau diese Überraschung
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
DEFINITION
Information theory measures information through uncertainty: the less likely an event, the more information it carries when it happens, and Shannon entropy is the average information of a random variable. Cross-entropy measures how much information it takes to encode data from one distribution using another, KL divergence is the excess, and perplexity is the language-model equivalent of cross-entropy.
Intuition
"The sun will rise tomorrow" carries almost no information because you expected it; "there will be a total solar eclipse tomorrow" carries a great deal. How much information a message carries depends on how much it surprises you. Training a language model is the same goal: make it as unsurprised as possible by the next piece of real text. The loss it reports is the average of exactly that surprise.
Shannon’s model of communication: information content sets the limit of compression, and entropy is that limit
Bernoulli entropy and cross-entropy: when the model encodes with a uniform q=0.5 the cross-entropy stays at 1 bit, and its gap above the true entropy H(p) is exactly the KL divergence (click the legend to toggle)
- True entropy H(p)
- Cross-entropy H(p, q=0.5)
Funktionsweise
- 01
Information content: the more surprising, the more informative
Define the information of an event as −log p: the smaller the probability p, the larger the information; the information of two independent events adds up, mirroring how their probabilities multiply. The logarithm is exactly what makes "information adds" match "probability multiplies", and with base 2 the unit is the bit.
- 02
Entropy: an average measure of uncertainty
Shannon entropy H(X) = −Σ p log p is the weighted average of the information of all values, measuring how hard the random variable is to guess. A fair coin has 1 bit of entropy, the hardest case; nearly certain events approach 0. Over a fixed value set, the more uniform the distribution, the larger the entropy.
- 03
Cross-entropy: the cost of encoding reality with the wrong model
Cross-entropy H(P,Q) = −Σ p(x) log q(x). Here the "true distribution" is P (the data) and the model supplies Q. It equals the true entropy H(P) plus the extra cost of using the wrong distribution. Training a language model is minimising cross-entropy — that is, minimising the model’s surprise at real text.
- 04
KL divergence and perplexity: spelling out the cost and the hesitation
KL divergence D(P‖Q) = H(P,Q) − H(P) is the part of the cross-entropy beyond the true entropy; it is non-negative and asymmetric, so it is not a distance. Perplexity = 2^{H(P,Q)} (or e^{H} with natural logs) reads as "how many equally likely options the model is hesitating among per step". In language-model papers, perplexity is often the more intuitive number.
Kernformel
H(P, Q) = − Σₓ P(x) log Q(x)Perplexity of language models on WikiText-103: lower means the model is less surprised by real text (indicative magnitudes from public reports)
Anwendungsfelder
- The training objective of language models: cross-entropy loss is negative log-likelihood
- Evaluation metric: perplexity compares how well different models model text
- Data compression: entropy is the theoretical limit on compression and guides the design of entropy coders
- Decision trees split on "information gain" (a drop in entropy); variational inference minimises KL to approximate a posterior
Häufige Missverständnisse
- Cross-entropy punishes confident mistakes very steeply: the more probability the model stakes on a wrong answer, the larger the loss. Conversely, underrating the correct answer is also costly — which is exactly what forces the model to be honest.
- KL divergence is asymmetric, D(P‖Q) ≠ D(Q‖P), so it cannot be used as a distance; optimising the two directions gives very different results (one tends to cover, the other to lock onto modes).
- Perplexity is comparable only under the same vocabulary and tokenisation. Comparing perplexity across models or tokenisers is badly misleading — finer tokenisation tends to lower perplexity without meaning a better model.
Schlüsselbegriffe
- Self-information
- The information of a single event, −log p
- Shannon entropy
- The average information, or uncertainty, of a random variable
- Cross-entropy
- The information needed to encode data from P using distribution Q
- KL divergence
- Cross-entropy minus true entropy; non-negative and asymmetric
- Perplexity
- The exponential of the cross-entropy; the effective number of options the model hesitates among per step