Maximum Likelihood Estimation finds the "most likely" distribution parameters
Given a set of data, search for the mean and variance that maximize the likelihood function, and compare whether the results are consistent with np.mean / np.std.
After completing this case, you will understand:The parameters that "maximize the probability of the data occurring" are exactly the sample mean and sample standard deviation.
Real-life Introduction
Guess the average weight of fish in a pond
You don't know the average weight of all fish in the pond. But you caught 200 fish and weighed them. You naturally wonder: What kind of "average weight" would make the event "exactly weighing these weights of these 200 fish" most likely to happen?
Intuition tells you that the average weight of the 200 fish is a very good estimate. Maximum Likelihood Estimation (MLE) mathematizes this intuition: find the parameter value that maximizes the probability of observing this set of data.
Mathematical Definition
Log-likelihood of the Gaussian Distribution
\[ \log L(\mu, \sigma) = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(x_i - \mu)^2 \]Analytical Solution of MLE
\[ \hat{\mu}_{\text{MLE}} = \frac{1}{n}\sum_{i=1}^{n} x_i = \bar{x}, \quad \hat{\sigma}_{\text{MLE}} = \sqrt{\frac{1}{n}\sum (x_i - \bar{x})^2} \]This is not a coincidence. The analytical solution of MLE for a Gaussian distribution is exactly the sample mean and sample standard deviation.
Python Hands-on Practice
Example
np.random.seed(3)
true_mu, true_sigma = 5.0, 2.0
data = np.random.normal(true_mu, true_sigma, size=200)
# Gaussian log-likelihood function
def log_likelihood(mu, sigma, data):
n = len(data)
return (-n/2) * np.log(2 * np.pi * sigma**2) \
- np.sum((data - mu)**2) / (2 * sigma**2)
# Brute-force grid search for MLE
mu_grid = np.linspace(3, 7, 100)
sigma_grid = np.linspace(0.5, 4, 100)
best_ll = -np.inf
best_mu, best_sigma = None, None
for mu in mu_grid:
for sigma in sigma_grid:
ll = log_likelihood(mu, sigma, data)
if ll > best_ll:
best_ll = ll
best_mu, best_sigma = mu, sigma
print("=" * 55)
print("EXAMPLE Maximum Likelihood Estimation (MLE) result")
print("=" * 55)
print(f"Grid search MLE: mu={best_mu:.4f}, sigma={best_sigma:.4f}")
print(f"Direct formula: mu={data.mean():.4f}, sigma={data.std():.4f}")
print(f"True generation parameters: mu={true_mu}, sigma={true_sigma}")
print(f"\nConclusion: MLE = np.mean/np.std — not a coincidence, but a mathematical theorem")
======================================================= EXAMPLE 最大似然估计 (MLE) 结果 ======================================================= 网格搜索 MLE: mu=4.9394, sigma=2.0480 公式直接算: mu=4.9394, sigma=2.0480 真实生成参数: mu=5.0, sigma=2.0 结论: MLE = np.mean/np.std —— 不是巧合,是数学定理
Application Scenarios in AI
| Scenario | Connection with MLE |
|---|---|
| Cross-entropy Loss | Minimizing cross-entropy in classification tasks = MLE of the categorical distribution |
| MSE Loss | Minimizing MSE in regression tasks = MLE assuming the error follows a Gaussian distribution |
| Parameter initialization | Initializing model parameters with data statistics—essentially the intuition of MLE. |