الإنتروبيا ونظرية المعلومات
كلما كانت الجملة أكثر مفاجأة حملت معلومات أكثر، ودالة الخسارة في نموذج اللغة تقيس هذه المفاجأة نفسها
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
التعريف
Information theory measures information through uncertainty: the less likely an event, the more information it carries when it happens, and Shannon entropy is the average information of a random variable. Cross-entropy measures how much information it takes to encode data from one distribution using another, KL divergence is the excess, and perplexity is the language-model equivalent of cross-entropy.
الحدس المباشر
"The sun will rise tomorrow" carries almost no information because you expected it; "there will be a total solar eclipse tomorrow" carries a great deal. How much information a message carries depends on how much it surprises you. Training a language model is the same goal: make it as unsurprised as possible by the next piece of real text. The loss it reports is the average of exactly that surprise.
Shannon’s model of communication: information content sets the limit of compression, and entropy is that limit
Bernoulli entropy and cross-entropy: when the model encodes with a uniform q=0.5 the cross-entropy stays at 1 bit, and its gap above the true entropy H(p) is exactly the KL divergence (click the legend to toggle)
- True entropy H(p)
- Cross-entropy H(p, q=0.5)
طريقة العمل
- 01
Information content: the more surprising, the more informative
Define the information of an event as −log p: the smaller the probability p, the larger the information; the information of two independent events adds up, mirroring how their probabilities multiply. The logarithm is exactly what makes "information adds" match "probability multiplies", and with base 2 the unit is the bit.
- 02
Entropy: an average measure of uncertainty
Shannon entropy H(X) = −Σ p log p is the weighted average of the information of all values, measuring how hard the random variable is to guess. A fair coin has 1 bit of entropy, the hardest case; nearly certain events approach 0. Over a fixed value set, the more uniform the distribution, the larger the entropy.
- 03
Cross-entropy: the cost of encoding reality with the wrong model
Cross-entropy H(P,Q) = −Σ p(x) log q(x). Here the "true distribution" is P (the data) and the model supplies Q. It equals the true entropy H(P) plus the extra cost of using the wrong distribution. Training a language model is minimising cross-entropy — that is, minimising the model’s surprise at real text.
- 04
KL divergence and perplexity: spelling out the cost and the hesitation
KL divergence D(P‖Q) = H(P,Q) − H(P) is the part of the cross-entropy beyond the true entropy; it is non-negative and asymmetric, so it is not a distance. Perplexity = 2^{H(P,Q)} (or e^{H} with natural logs) reads as "how many equally likely options the model is hesitating among per step". In language-model papers, perplexity is often the more intuitive number.
الصيغة الأساسية
H(P, Q) = − Σₓ P(x) log Q(x)Perplexity of language models on WikiText-103: lower means the model is less surprised by real text (indicative magnitudes from public reports)
مجالات الاستخدام
- The training objective of language models: cross-entropy loss is negative log-likelihood
- Evaluation metric: perplexity compares how well different models model text
- Data compression: entropy is the theoretical limit on compression and guides the design of entropy coders
- Decision trees split on "information gain" (a drop in entropy); variational inference minimises KL to approximate a posterior
مفاهيم خاطئة شائعة
- Cross-entropy punishes confident mistakes very steeply: the more probability the model stakes on a wrong answer, the larger the loss. Conversely, underrating the correct answer is also costly — which is exactly what forces the model to be honest.
- KL divergence is asymmetric, D(P‖Q) ≠ D(Q‖P), so it cannot be used as a distance; optimising the two directions gives very different results (one tends to cover, the other to lock onto modes).
- Perplexity is comparable only under the same vocabulary and tokenisation. Comparing perplexity across models or tokenisers is badly misleading — finer tokenisation tends to lower perplexity without meaning a better model.
مصطلحات أساسية
- Self-information
- The information of a single event, −log p
- Shannon entropy
- The average information, or uncertainty, of a random variable
- Cross-entropy
- The information needed to encode data from P using distribution Q
- KL divergence
- Cross-entropy minus true entropy; non-negative and asymmetric
- Perplexity
- The exponential of the cross-entropy; the effective number of options the model hesitates among per step