Maximum Likelihood Estimation finds the "most likely" distribution parameters

Given a set of data, search for the mean and variance that maximize the likelihood function, and compare whether the results are consistent with np.mean / np.std.

After completing this case, you will understand:The parameters that "maximize the probability of the data occurring" are exactly the sample mean and sample standard deviation.


Real-life Introduction

Guess the average weight of fish in a pond

You don't know the average weight of all fish in the pond. But you caught 200 fish and weighed them. You naturally wonder: What kind of "average weight" would make the event "exactly weighing these weights of these 200 fish" most likely to happen?

Intuition tells you that the average weight of the 200 fish is a very good estimate. Maximum Likelihood Estimation (MLE) mathematizes this intuition: find the parameter value that maximizes the probability of observing this set of data.


Mathematical Definition

Log-likelihood of the Gaussian Distribution

\[ \log L(\mu, \sigma) = -\frac{n}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{n}(x_i - \mu)^2 \]

Analytical Solution of MLE

\[ \hat{\mu}_{\text{MLE}} = \frac{1}{n}\sum_{i=1}^{n} x_i = \bar{x}, \quad \hat{\sigma}_{\text{MLE}} = \sqrt{\frac{1}{n}\sum (x_i - \bar{x})^2} \]

This is not a coincidence. The analytical solution of MLE for a Gaussian distribution is exactly the sample mean and sample standard deviation.


Python Hands-on Practice

Example

import numpy as np
np.random.seed(3)

true_mu, true_sigma = 5.0, 2.0
data = np.random.normal(true_mu, true_sigma, size=200)

# Gaussian log-likelihood function
def log_likelihood(mu, sigma, data):
    n = len(data)
    return (-n/2) * np.log(2 * np.pi * sigma**2) \
           - np.sum((data - mu)**2) / (2 * sigma**2)

# Brute-force grid search for MLE
mu_grid = np.linspace(3, 7, 100)
sigma_grid = np.linspace(0.5, 4, 100)
best_ll = -np.inf
best_mu, best_sigma = None, None

for mu in mu_grid:
    for sigma in sigma_grid:
        ll = log_likelihood(mu, sigma, data)
        if ll > best_ll:
            best_ll = ll
            best_mu, best_sigma = mu, sigma

print("=" * 55)
print("EXAMPLE Maximum Likelihood Estimation (MLE) result")
print("=" * 55)
print(f"Grid search MLE: mu={best_mu:.4f}, sigma={best_sigma:.4f}")
print(f"Direct formula: mu={data.mean():.4f}, sigma={data.std():.4f}")
print(f"True generation parameters: mu={true_mu}, sigma={true_sigma}")
print(f"\nConclusion: MLE = np.mean/np.std — not a coincidence, but a mathematical theorem")
=======================================================
EXAMPLE 最大似然估计 (MLE) 结果
=======================================================
网格搜索 MLE: mu=4.9394,  sigma=2.0480
公式直接算:   mu=4.9394,  sigma=2.0480
真实生成参数: mu=5.0, sigma=2.0

结论: MLE = np.mean/np.std —— 不是巧合,是数学定理

Application Scenarios in AI

ScenarioConnection with MLE
Cross-entropy LossMinimizing cross-entropy in classification tasks = MLE of the categorical distribution
MSE LossMinimizing MSE in regression tasks = MLE assuming the error follows a Gaussian distribution
Parameter initializationInitializing model parameters with data statistics—essentially the intuition of MLE.
Other extensions