Maximum Likelihood Estimation and Maximum A Posteriori Estimation
MLE and MAP are two core methods of statistical inference.The essence of training neural networks ≈ maximum likelihood estimation.
Concept Analysis
Suppose you have a set of observed data, and you believe these data come from a known type of probability distribution (e.g., normal distribution), but you don't know the parameters of this distribution (mean μ and standard deviation σ).
MLE and MAP answer the same question: given the data, what is the most reasonable parameter value?
The difference is that MLE only looks at the data itself, while MAP also references your prior knowledge about the parameters.
Likelihood vs Probability: Two Ways to View the Same Thing
Probability P(data|parameters): parameters are fixed, asking "Given these parameters, how likely is it to observe these data?"
Likelihood L(parameters|data): data are fixed, asking "Which parameter value is most likely to produce these data?"
The value of the likelihood function itself is not normalized (the sum does not need to be 1); its significance lies incomparing the relative plausibility of different parameter values。
MLE: Which parameters are most likely to generate these data?
\[ \hat{\theta}_{\text{MLE}} = \arg\max_\theta P(D | \theta) \]Looking only at the data, choose the parameters that most likely produce the observed data.
Why take the log?
- Product becomes sum: numerically more stable (prevents underflow)
- Differentiation is easier: logarithm turns multiplication into addition
- Does not change extrema: log is a monotonically increasing function
MAP: Incorporating Prior Knowledge
\[ \hat{\theta}_{\text{MAP}} = \arg\max_\theta P(\theta | D) = \arg\max_\theta P(D | \theta) \cdot P(\theta) \]MAP = MLE + prior P(θ). When the prior is a uniform distribution, MAP = MLE.
MLE only looks at data, MAP also looks at the prior. When data is large, the influence of the prior is diluted, and MAP approaches MLE.
When data is small, the prior acts as "regularization" — L2 regularization ≡ MAP with a Gaussian prior.
Real-life Example
Estimating average daily sales
Online store daily sales for 10 consecutive days: [23, 25, 22, 24, 26, 23, 25, 24, 22, 24].
Assume sales follow a normal distribution. The MLE estimate of the mean = the average of these numbers ≈ 23.8.
Intuitively you would do this too — MLE just turns intuition into mathematics.
Hands-on Practice with Python
Example
np.random.seed(42)
data = np.random.normal(5.0, 2.0, 100) # True μ=5, σ=2
# MLE: μ = sample mean
mu_mle = np.mean(data)
print(f"MLE μ = {mu_mle:.3f} (true=5.0)")
# MAP: prior μ ~ N(0, 1), shrinks toward the prior
prior_mu, prior_var = 0.0, 1.0
n = len(data)
mu_map = (prior_mu/prior_var + n*mu_mle/np.std(data)**2) / (1/prior_var + n/np.std(data)**2)
print(f"MAP μ = {mu_map:.3f} (compromise between MLE and prior 0)")
Running output:
MLE μ = 4.968 (真实=5.0) MAP μ = 4.776 (在MLE和先验0之间折中)
Application Scenarios in AI
Loss Function Design = MLE
Minimizing cross-entropy loss ≈ maximum likelihood estimation under a multinomial distribution assumption. Minimizing MSE ≈ MLE under a normal error distribution assumption. The loss function is not chosen arbitrarily — each loss function corresponds to an implicit probability distribution assumption.
L2 Regularization = MAP with Gaussian Prior
Adding \( \lambda\|w\|^2 \) to the loss is equivalent to assuming the prior of weights w is \( N(0, 1/\lambda) \). MAP compromises between data likelihood and prior — when data is scarce, the prior dominates (weights tend to 0); when data is abundant, the likelihood dominates.
L1 Regularization = MAP with Laplace Prior
Adding \( \lambda\|w\|_1 \) to the loss is equivalent to assuming the prior of weights w is \( Laplace(0, 1/\lambda) \) distribution. The Laplace distribution is sharp at 0 — MAP tends to produce sparse solutions (many weights exactly 0), enabling automatic feature selection.
Bayesian Neural Networks
The weights of an ordinary neural network are deterministic point estimates (MLE/MAP). A Bayesian neural network assigns a probability distribution (posterior) to each weight, and the output is also a probability distribution — thus providing natural uncertainty estimation. During prediction, you can sample the weights multiple times and look at the variance of the outputs to judge the model's confidence in its own predictions.
Other extensions