Discrete Random Variables -- From Dice Rolling to Probability Mass Function
A random variable formalizes "random events" as a mathematical tool.A discrete random variable takes a finite or countable number of values.
Concept Analysis
Random Variable
Random variable = a variable whose value is uncertain, with a probability associated with each value.
PMF (Probability Mass Function)
\[ P(X = x_k) = p_k, \quad \sum_k p_k = 1 \]Bernoulli Distribution
One trial, two outcomes: P(success)=p, P(failure)=1-p. Mathematical modeling for binary classification problems.
Binomial Distribution
\[ P(X=k) = \binom{n}{k} p^k (1-p)^{n-k} \]The probability of k successes in n independent Bernoulli trials.
Real-life Examples
Roll a die: X is the face value, P(X=1)=1/6, P(X=2)=1/6, ...
Toss a coin 10 times; the number of heads X follows a binomial distribution Binomial(10, 0.5), and the most likely value is 5.
Python Hands-on Practice
Example
# Rolling Dice PMF
rolls = np.random.randint(1, 7, size=10000)
for v in range(1, 7):
print(f"Value {v}: {np.mean(rolls==v):.4f}")
# Bernoulli
print(f"\n伯努利 p=0.3, 1000times试验, success率={np.random.binomial(1,0.3,1000).mean():.3f}")
# Binomial distribution: Binomial(10, 0.3)
samples = np.random.binomial(10, 0.3, 10000)
k_best = np.bincount(samples).argmax()
print(f"Binomial distribution n=10,p=0.3: most likely k={k_best} (theoretical=np=3)")
Output:
点数1: 0.1628 点数2: 0.1661 点数3: 0.1686 点数4: 0.1686 点数5: 0.1637 点数6: 0.1702 伯努利 p=0.3, 1000次试验, 成功率=0.295 二项分布 n=10,p=0.3: 最可能 k=3 (理论=np=3)
Binomial Distribution PMF Interactive Chart
Probability mass function (PMF) of the binomial distribution Binomial(n=10, p=0.3). Adjust the slider to observe the distribution shape under different p values:
Application Scenarios in AI
Binary classification labels = Bernoulli distribution
Logistic regression and binary classifiers model the label y ∈ {0,1} as a Bernoulli distribution: P(y=1|x) = \sigma(wx+b), P(y=0|x) = 1-\sigma(wx+b). The binary cross-entropy loss is the negative log-likelihood under this assumption. This is why the output layer for binary classification tasks uses Sigmoid + BCE.
Multi-class classification = Categorical distribution (a single trial of a multinomial distribution)
In multi-class classification, labels are one-hot vectors, and the probabilities of each class output by Softmax sum to 1. This is equivalent to a single trial of an n-sided "die" with different probabilities for each face. Under this assumption, the cross-entropy loss equals the negative log-likelihood.
Dropout regularization = Bernoulli trial
During training, each neuron is retained with probability p (a Bernoulli trial) and dropped with probability 1-p. This is equivalent to training a different subnetwork at each iteration, achieving an ensemble learning effect. At test time, all neurons are active but the outputs are scaled by p as compensation.
Exploration in reinforcement learning = Sampling from a discrete distribution
DQN often uses the ε-greedy strategy: with probability ε, choose a random action (Bernoulli), and with probability 1-ε, choose the optimal action. This is the simplest exploration strategy, ensuring the agent does not get stuck in a suboptimal policy forever.