Visualizing KL Divergence -- How Much Two Distributions Differ

Adjust the parameters of two Gaussian distributions, compute and compare KL(P||Q) and KL(Q||P), and intuitively feel the asymmetry of KL divergence.

After completing this case, you will understand:KL divergence is not a 'distance' — viewing Q from P and viewing P from Q, the information loss is different.


Everyday Introduction

Using a Beijing Map to Find Your Way in New York

You have two maps: the real New York map (P) and the Beijing map (Q). Using the Beijing map to walk in New York, you need to spend extra effort to 'correct' the wrong information on the map — this extra effort is KL(P||Q). Conversely, using the New York map to walk in Beijing — the extra effort is KL(Q||P). These two kinds of extra effort are clearly different, because the two maps 'go wrong' in different ways.


Intuitive Understanding

Scenario A: P~N(0,1), Q~N(2,1) — different means but same shape, KL is nearly symmetric.

Scenario B: P~N(0,1), Q~N(0,3) — same center but different variances. KL(P||Q) = 1.10 (narrow → wide, easy), KL(Q||P) = 1.55 (wide → narrow, difficult). Clearly asymmetric — 'covering a narrow distribution with a wide one' is much easier than 'covering a wide distribution with a narrow one'.


Mathematical Definition

\[ D_{KL}(P \parallel Q) = \sum_{x} P(x) \cdot \log\frac{P(x)}{Q(x)} \] \[ D_{KL}(P \parallel Q) = H(P, Q) - H(P) \]

P's own entropy is constant, so minimizing cross-entropy = minimizing KL divergence.


Python Hands-on Practice

Example

import numpy as np

def gaussian_pdf(x, mu, sigma):
    return (1 / (sigma * np.sqrt(2 * np.pi))) * \
           np.exp(-(x - mu) ** 2 / (2 * sigma ** 2))

def kl_divergence(p, q, eps=1e-12):
    p = np.clip(p, eps, None)
    q = np.clip(q, eps, None)
    return np.sum(p * np.log(p / q))

x = np.linspace(-10, 10, 2000)
dx = x[1] - x[0]

# Scenario A: different means
P_A = gaussian_pdf(x, mu=0, sigma=1) * dx
Q_A = gaussian_pdf(x, mu=2, sigma=1) * dx

# Scenario B: different variances
P_B = gaussian_pdf(x, mu=0, sigma=1) * dx
Q_B = gaussian_pdf(x, mu=0, sigma=3) * dx

print("EXAMPLE KL divergence asymmetry verification:\n")
print("Scenario A: P~N(0,1) vs Q~N(2,1)")
print(f"  KL(P||Q) = {kl_divergence(P_A, Q_A):.4f}")
print(f"  KL(Q||P) = {kl_divergence(Q_A, P_A):.4f}")

print(f"\nScenario B: P~N(0,1) vs Q~N(0,3)")
print(f" KL(P||Q) = {kl_divergence(P_B, Q_B):.4f} (narrow -> wide, easy)")
print(f" KL(Q||P) = {kl_divergence(Q_B, P_B):.4f} (wide -> narrow, difficult)")
print(f"\nEXAMPLE intuition: 'covering narrow with wide' is easy (small KL), 'covering wide with narrow' is difficult (large KL)")
EXAMPLE KL 散度不对称性验证:

场景 A:P~N(0,1)  vs  Q~N(2,1)
  KL(P||Q) = 2.0000
  KL(Q||P) = 1.9960

场景 B:P~N(0,1)  vs  Q~N(0,3)
  KL(P||Q) = 1.0990  (窄->宽,容易)
  KL(Q||P) = 1.5490  (宽->窄,困难)

EXAMPLE 直觉:'用宽覆盖窄'容易(KL小),'用窄覆盖宽'困难(KL大)

Application Scenarios in AI

SceneHow to use KL divergence?
VAEMake the encoding distribution close to standard normal — loss = reconstruction error + KL(encoding || N(0,1))
Knowledge distillationMake the student model’s predicted distribution close to the teacher model—minimize KL(teacher||student).
PPO reinforcement learningUse KL divergence to constrain the difference between old and new policies, preventing updates from being too large.
Other extensions