Probability Basics -- Sample Space, Events and Conditional Probability
Probability theory is a core tool for AI to handle uncertainty.From classifier outputs to generative model sampling, probability is everywhere.
Concept Analysis
Sample Space
The set of all possible outcomes
Rolling a die: {1,2,3,4,5,6}
Event
A subset of the sample space
Even = {2,4,6}
Probability P(A)
The likelihood of A occurring, [0,1]
0 = impossible, 1 = certain
Conditional Probability: Updating Judgment Based on Partial Information
\[ P(A|B) = \frac{P(A \cap B)}{P(B)} \]Intuition: narrow the sample space to the set where B occurs, and see what proportion of it is occupied by A.
Law of Total Probability: Summarizing Across Cases
\[ P(A) = \sum_i P(A|B_i) P(B_i) \]Everyday Examples
Law of Large Numbers
Flip a coin 10 times, you might get 4 heads and 6 tails. Flip it 10,000 times, the proportion of heads will definitely be very close to 0.5.
The more trials, the more stable the frequency—this is the foundation of probability theory.
Python Hands-On Practice
Example
# Law of Large Numbers Verification
for n in [10, 100, 1000, 10000]:
tosses = np.random.choice(['H','T'], size=n)
freq = np.mean(tosses == 'H')
print(f"EXAMPLE flipped {n:5d} times, head frequency={freq:.4f}")
# Monty Hall problem: switching wins 2/3
def monty_hall(switch=True, trials=10000):
wins = 0
for _ in range(trials):
car = np.random.randint(0, 3)
choice = np.random.randint(0, 3)
revealed = np.random.choice([d for d in range(3) if d != car and d != choice])
if switch:
choice = [d for d in range(3) if d != choice and d != revealed][0]
wins += (choice == car)
return wins/trials
print(f"\n"Monty Hall: no switch={monty_hall(False):.3f}, switch={monty_hall(True):.3f} (theoretical: 1/3, 2/3)")
Output:
EXAMPLE 掷 10次, 正面频率=0.8000 EXAMPLE 掷 100次, 正面频率=0.5500 EXAMPLE 掷 1000次, 正面频率=0.4870 EXAMPLE 掷10000次, 正面频率=0.4987 蒙提霍尔: 不换=0.329, 换=0.663 (理论: 1/3, 2/3)
Application Scenarios in AI
Classifier output = probability distribution
Softmax outputs the probability for each class, and the sum of probabilities across all classes is 1. The model doesn't just tell you "this is a cat," it also tells you "80% probability it's a cat, 15% a dog, 5% something else." This kind of probabilistic output is crucial for risk assessment and confidence calibration.
Generative model = sampling from a probability distribution
VAE samples from a learned latent distribution to generate new images, GAN samples from a random noise distribution to generate realistic images, and diffusion models start from pure noise and gradually denoise to generate high-definition images. The core operation of all these generation processes is "sampling from some probability distribution."
Stochastic Policy in Reinforcement Learning
In policy gradient methods, the agent's actions are not deterministic, but are sampled from a probability distribution output by the policy network (e.g., 70% left, 30% right). This randomness ensures exploration—if you always choose the highest-probability action, you may never discover a better policy.
Other extensions