Cross-Entropy and KL Divergence -- Measuring the Difference Between Distributions
Cross-entropy is the most commonly used loss function for classification tasks.KL divergence measures how dissimilar two probability distributions are.
Concept Analysis
In the previous chapter, we learned about entropy — measuring the disorder of a single distribution itself. The two concepts in this chapter measurethe difference between two distributions, which directly determines the design of loss functions in classification tasks.
Cross-Entropy
\[ H(p, q) = -\sum_x p(x) \log q(x) \]Meaning: Using predicted distribution q to encode data from the true distribution p, how much information is spent on average.
KL Divergence (Relative Entropy)
\[ D_{KL}(p \| q) = \sum_x p(x) \log \frac{p(x)}{q(x)} = \underbrace{H(p,q)}_{\text{cross-entropy}} - \underbrace{H(p)}_{\text{true entropy}} \]Meaning: The extra information wasted after q replaces p. KL ≥ 0, and when it equals 0, p = q.
KL divergenceis asymmetric: D_KL(p||q) ≠ D_KL(q||p), so it is not a true "distance".
When the true distribution p is one-hot (deterministic), H(p) = 0, and cross-entropy equals KL divergence.
Why Does Classification Use Cross-Entropy Instead of MSE?
The gradient of cross-entropy + Softmax = \( \hat{y} - y \) (prediction - true), which is simple and efficient.
The gradient of MSE + Softmax approaches 0 when the prediction is confident — vanishing gradient, making learning impossible.
Real-Life Examples
Codebook design
You write letters using a set of codes (short codes represent common characters, long codes represent rare characters).
But if your codes are designed based on English frequency while you are writing Chinese—the coding efficiency is low.
Using the wrong distribution q to encode data from the true distribution p wastes extra bits = KL divergence.
Python Hands-On Practice
Example
def cross_entropy(p, q):
return -np.sum(p * np.log(q + 1e-10))
p_true = np.array([1.0, 0.0, 0.0]) # True: class 0
for q, desc in [([0.9,0.05,0.05],"Good prediction"),
([0.4,0.3,0.3],"Ambiguous"),
([0.1,0.8,0.1],"Wrong")]:
print(f"{desc}: cross-entropy={cross_entropy(p_true, q):.4f}")
# Compare gradients of CE and MSE
print("\n"Very bad prediction (y_pred=0.01, true=1): CE gradient=-100, MSE gradient=-1.98")
print("CE gives a large gradient when predictions are wrong → learns quickly")
Running output:
好预测: 交叉熵=0.1054 模糊: 交叉熵=0.9163 错误: 交叉熵=2.3026 预测很差(y_pred=0.01, 真=1): CE梯度=-100, MSE梯度=-1.98 CE在预测错误时给出大梯度 → 学得快
Cross-Entropy vs MSE Gradient Comparison
Application Scenarios in AI
Standard for Classification Tasks: Cross-Entropy Loss
PyTorch's nn.CrossEntropyLoss and TensorFlow's categorical_crossentropy are the default loss functions for classification tasks. They internally integrate Softmax + negative log-likelihood—a single loss function completes both activation and loss computation, and the gradient form is extremely simple: \( \nabla = \hat{y} - y \).
Why Not Use MSE for Classification? Gradient Properties Determine Efficiency
When the model predicts incorrectly (e.g., the true class probability should be close to 1 but is predicted as 0.01), the cross-entropy gradient is about -100—giving an extremely strong correction signal. MSE, on the other hand, has a gradient of only -1.98 in the same situation—it barely learns. This explains why MSE converges extremely slowly or even fails to converge in classification tasks.
KL Divergence in Knowledge Distillation
Knowledge distillation: use the Softmax output of a large model (teacher) as soft labels to train a small model (student). The student must not only match the true labels but also minimize the divergence between its own output distribution and the teacher's via KL divergence. Soft labels contain similarity information between classes (e.g., the output probabilities for "cat" and "tiger" are both higher than for "car"), which is dark knowledge that one-hot labels do not have.
KL Divergence Regularization Term in VAE
VAE's loss function = reconstruction error + KL(N(μ,σ²) || N(0,I)). The KL term constrains the latent distribution output by the encoder so that it does not deviate too far from the standard normal for each sample—this acts as a regularizer, ensuring that the latent space is continuous and smooth, and samples from different classes can transition smoothly in the latent space.
KL Constraint in RLHF
Models such as ChatGPT, when fine-tuned with RLHF (Reinforcement Learning from Human Feedback), add a KL divergence constraint to the optimization objective: do not let the current policy deviate too far from the initial SFT model. This prevents the model from outputting strange text to cater to the reward model, maintaining generation quality.
Other Extensions