Visualization of the Central Limit Theorem
Repeatedly sample from a uniform distribution that is completely non-normal, calculate the mean each time, and observe how the distribution of means approaches a bell shape.
After completing this case, you will understand:Why the Gaussian distribution is ubiquitous in AI — it is the inevitable result of the superposition of many independent factors.
Life Introduction
Why is the class average height always bell-shaped?
Measure the heights of 50 people in the class. An individual's height is influenced by hundreds of factors: genetics, nutrition, exercise, sleep... Each factor has a large or small impact, positive or negative.
When many independent small factors are stacked together, the final distribution is bell-shaped — this is a mathematical theorem, not a coincidence.The central limit theorem states: the sum (or mean) of a large number of independent random variables tends toward a normal distribution.
Intuitive Understanding
We sample from a uniform distribution U(0,1) — this distribution is completely flat and has nothing to do with a bell shape. But when we take multiple samples (e.g., 30) each time to compute one mean, and repeat this operation 5,000 times — the histogram of these 5,000 means begins to look bell-shaped. The larger n is, the closer it approaches a perfect normal distribution.
Mathematical Definition
\[ \frac{\bar{X}_n - \mu}{\sigma / \sqrt{n}} \xrightarrow{d} \mathcal{N}(0, 1) \]Where \(\bar{X}_n = \frac{1}{n}\sum X_i\). The standard deviation of the mean shrinks at a rate of \(1/\sqrt{n}\).
Python Hands-on Practice
Example
np.random.seed(2)
def sample_raw(size):
return np.random.uniform(0, 1, size=size)
def clt_experiment(n, repeat=5000):
means = [sample_raw(n).mean() for _ in range(repeat)]
return np.array(means)
sample_sizes = [1, 2, 5, 30]
results = {n: clt_experiment(n) for n in sample_sizes}
print("EXAMPLE Central Limit Theorem Verification:\n")
print("Original distribution: Uniform U(0,1) (completely flat)")
print("Theory: standard deviation of mean = 1/sqrt(12n)\n")
print(f"{'n':<6} {'Mean SD':<12} {'Theory':<12} {'Match'}")
print("-" * 42)
for n, means in results.items():
theo = 1 / np.sqrt(12 * n)
ok = "Yes" if abs(means.std() - theo) < 0.01 else "No"
print(f"{n:<6} {means.std():<12.4f} {theo:<12.4f} {ok}")
# Normality test for n=30
m30 = results[30]
skew = np.mean(((m30-m30.mean())/m30.std())**3)
kurt = np.mean(((m30-m30.mean())/m30.std())**4) - 3
print(f"\nEXAMPLE n=30: skewness={skew:.3f} (should be ~0), excess kurtosis={kurt:.3f} (should be ~0)")
EXAMPLE 中心极限定理验证: 原始分布: 均匀分布 U(0,1)(完全平的) 理论: 均值的标准差 = 1/sqrt(12n) n 均值SD 理论值 匹配 ------------------------------------------ 1 0.2891 0.2887 Yes 2 0.2043 0.2041 Yes 5 0.1292 0.1291 Yes 30 0.0526 0.0527 Yes EXAMPLE n=30 时:偏度=-0.023 (应~0), 超峰度=-0.005 (应~0)
Application Scenarios in AI
| Scenario | Connection to CLT |
|---|---|
| Batch Normalization | Assume that the activation values within each mini-batch are approximately normally distributed, then subtract the mean and divide by the standard deviation. |
| Weight initialization | Xavier/He initialization samples initial weights from a normal distribution |
| Error modeling | In regression problems, it is assumed that errors follow a normal distribution—errors are the result of the superposition of multiple unmodeled factors. |