Information Content and Entropy -- Quantifying Uncertainty
Information theory provides AI with mathematical tools for quantifying "information" and "uncertainty".
The less likely an event is, the more information it carries once it occurs.
Concept Analysis
Information theory seeks to answer a fundamental question:How do we quantify the amount of "information" contained in a message?
Key insight: Information content is inversely proportional to the probability of an event — the less likely an event is, the more information it tells you once it occurs.
Why is information content defined as -log P(x)?
This definition encompasses three most natural intuitions:
- Inverse relationship: The smaller P(x) is (the more surprising), the larger I(x) should be. The negative sign ensures this.
- Additivity: The information content of two independent events occurring together equals the sum of their individual information contents. The logarithm turns the product of probabilities into addition of information content: \( -\log(p_1 \cdot p_2) = -\log p_1 + (-\log p_2) \)
- Certain events are zero: When a certain event has P=1, \( -\log 1 = 0 \) — known facts bring no information.
Information Content (Self-Information)
\[ I(x) = -\log P(x) \]Base 2 → unit bits, base e → unit nats. In AI, the natural logarithm is commonly used (convenient for differentiation).
Entropy: Expected Value of Information Content
\[ H(X) = -\sum_x P(x) \log P(x) \]Entropy measures the uncertainty of the entire probability distribution.
- Certain events (P=1): H = 0 (no uncertainty at all)
- Uniform distribution: H reaches its maximum (most uncertain)
Higher entropy = harder to predict. This provides the theoretical foundation for decision tree splitting criteria (information gain) and exploration strategies in reinforcement learning.
Real-Life Examples
News Value
"Tomorrow the sun will rise" → information content ≈ 0 (you already knew that).
「明天Yes 8 levelground震」→ extremely high information content (an extremely rare event happened).
The more surprising the news, the greater the information content—this is the intuition behind I(x) = -log P(x).
Python Hands-On Practice
Example
def entropy(probs):
probs = np.array(probs)
probs = probs[probs > 0]
return -np.sum(probs * np.log(probs))
print("Certain distribution [1,0,0]:", entropy([1,0,0])) # 0
print("2-class uniform [0.5,0.5]:", entropy([0.5,0.5])) # 0.693
print("10-class uniform:", entropy([0.1]*10)) # 2.303
print("Biased [0.9,0.05,0.05]:", entropy([0.9,0.05,0.05])) # 0.394
Run output:
确定分布 [1,0,0]: 0.0 2类均匀 [0.5,0.5]: 0.6931471805599453 10类均匀: 2.302585092994046 有偏 [0.9,0.05,0.05]: 0.3944479480436973
Binary Classification Entropy Curve
Application Scenarios in AI
Decision Tree Splitting Criterion: Information Gain
When a decision tree chooses a splitting feature, it computes the difference in entropy before and after the split — information gain = H(before split) - H(after split). It selects the feature and split point with the greatest information gain. The C4.5 and ID3 algorithms are both based on this principle.
Cross-Entropy Loss = Minimizing the Difference Between Predicted Distribution and True Distribution
In classification tasks, the entropy of the true distribution p (one-hot labels) is H(p)=0, and the cross-entropy H(p,q) equals KL(p||q). Minimizing cross-entropy = making the predicted distribution q as close as possible to the true distribution p. The next chapter will elaborate.
Entropy Regularization in Reinforcement Learning
The SAC (Soft Actor-Critic) algorithm adds a policy entropy term to the optimization objective: maximize reward + \( \alpha \cdot H(\pi) \). This encourages the policy to maintain a certain amount of randomness (not being too certain), promotes exploration, and prevents premature convergence to suboptimal policies. \(\alpha\) is the entropy regularization coefficient, controlling the balance between exploration and exploitation.
Entropy in Model Calibration
A good classifier not only has high accuracy, but the entropy of its predicted distribution should also reflect true uncertainty. When the model encounters out-of-distribution (OOD) samples, ideally the entropy of the Softmax output should be high (close to a uniform distribution), indicating "I am uncertain." If the entropy is very low but the prediction is wrong, the model is overconfident.
Other extensions