Bayes' Theorem -- Using New Evidence to Update Your Judgment

Bayes' Theorem is the cornerstone of probabilistic reasoning.Prior belief + new evidence → updated belief.


Concept Analysis

Prior P(H)
Belief before seeing evidence
×
Likelihood P(E|H)
Credibility of evidence
→
Posterior P(H|E)
Updated belief
\[ P(H|E) = \frac{P(E|H) \cdot P(H)}{P(E)} \]

P(E) is the normalization factor: \( P(E) = P(E|H)P(H) + P(E|\text{not }H)P(\text{not }H) \).

Core idea:Continuously update our judgment with new data. This is the basic pattern of all Bayesian methods.


Everyday Example

The False Positive Trap in Disease Testing

A certain disease has a prevalence of 0.1% and a test accuracy of 99%. If you test positive—the actual probability of having the disease is only about 9%.

Why? Because the disease is extremely rare, the absolute number of false positives (1% of healthy people misdiagnosed) far exceeds true positives.

Bayes' Theorem reveals this counterintuitive conclusion:When the prior probability is extremely low, even a highly accurate test means a positive result is likely a false positive.。


Python Hands-on Practice

Example

# Bayesian analysis of disease testing
P_H = 0.001           # Prior: prevalence rate 0.1%
P_E_given_H = 0.99    # Sensitivity 99%
P_E_given_notH = 0.01 # False positive rate 1%

P_E = P_E_given_H * P_H + P_E_given_notH * (1 - P_H)
P_H_given_E = P_E_given_H * P_H / P_E

print("=== EXAMPLE Disease Testing ===")
print(f"Prior prevalence: {P_H:.1%}")
print(f"After testing positive, actual disease probability: {P_H_given_E:.1%}")
print(f"Conclusion: With extremely low base probability, false positives far exceed true positives")

Output:

=== EXAMPLE 疾病检测 ===
先验患病率: 0.1%
检测阳性后,实际患病概率: 9.0%
结论:极低基础概率下,假阳性远超真阳性

Application Scenarios in AI

Naive Bayes Classifier

Assuming features are conditionally independent, use Bayes' theorem to compute the posterior probability P(category|features), and select the category with the largest posterior. It works surprisingly well in text classification (spam filtering, sentiment analysis)—although the "naive" independence assumption often does not hold in practice, classification accuracy is still very high.

Bayesian Optimization for Hyperparameter Tuning

When selecting hyperparameters such as learning rate and number of layers, each training run is time-consuming. Bayesian optimization uses a Gaussian process as a prior, updates the posterior distribution based on evaluated hyperparameter combinations (evidence), and intelligently selects the next most promising set of hyperparameters. It is much more efficient than grid search and random search.

Variational Autoencoder (VAE)

The encoder of a VAE outputs the mean μ and variance σ² of the latent distribution, then uses KL divergence to constrain this distribution to be close to the standard normal prior N(0,1). The "prior → posterior" framework here is Bayes' theorem—the prior is N(0,1), the likelihood comes from the decoder's reconstruction error, and the posterior is the distribution output by the encoder.


Other extensions