Naive Bayes Spam Classifier
Implement a simple Chinese spam classifier using the most basic conditional probability formula, without calling sklearn.
After completing this case, you will understand:Bayes' theorem + the "word independence" assumption = a usable text classifier.
Everyday Introduction
Email spam filtering
Your mailbox receives 100 emails every day. How can the program automatically determine whether a new email is spam? The program's approach: count, in historical emails, the probability of words like "free" and "winning" appearing in spam, and the probability of words like "meeting" and "project" appearing in normal emails.
When a new email arrives, calculate: the probability of these words appearing together under the spam assumption vs. under the normal email assumption. Whichever is larger wins the classification.
Intuitive Understanding
The key simplification of Naive Bayes — the "naive assumption": it assumes that the occurrence of each word is independent of the others. Although this assumption does not hold in reality ("free" and "prize" often appear together), its practical performance is surprisingly good.
| Techniques | Approach | Problem Solved |
|---|---|---|
| Logarithmic Probability | Use log to convert multiplication into addition | Multiplying many small decimals can become 0 (numerical underflow) |
| Laplace Smoothing | Add 1 to each word count instead of 0 | Words not seen in the training set cause probability to be 0 |
Python Hands-On Practice
Example
import math
# Training data (8 labeled Chinese SMS messages)
train_data = [
(Win Free Prizes Click Now, "spam"),
(Congratulations Winning Claim Cash Discount, "spam"),
(Limited Time Offer Click Link Free, "spam"),
(Loan Low Interest Apply Now Cash, "spam"),
(Meeting Arrangement Tomorrow Morning Discussion, "ham"),
(Project Progress Report Meeting Minutes, "ham"),
(Weekend Dinner Arrangement Everyone Attend, "ham"),
(Report Data Analysis Project Completion, "ham"),
]
# Count word frequencies for each category
class_word_counts = defaultdict(lambda: defaultdict(int))
class_doc_counts = defaultdict(int)
vocab = set()
for text, label in train_data:
class_doc_counts[label] += 1
for word in text.split():
class_word_counts[label][word] += 1
vocab.add(word)
vocab_size = len(vocab)
total_docs = len(train_data)
# Naive Bayes classification (log probability + Laplace smoothing)
def predict(text, alpha=1.0):
words = text.split()
scores = {}
for label in class_doc_counts:
log_prob = math.log(class_doc_counts[label] / total_docs)
total_words = sum(class_word_counts[label].values())
for word in words:
word_count = class_word_counts[label][word]
word_prob = (word_count + alpha) / (total_words + alpha * vocab_size)
log_prob += math.log(word_prob)
scores[label] = log_prob
return max(scores, key=scores.get), scores
# Test
test_emails = [
Claim Free Cash Prizes,
Tomorrow's meeting project report,
Time-limited discount link,
On the weekend, everyone discusses the report together.,
]
print(EXAMPLE Naive Bayes classification result:\n")
for email in test_emails:
label, scores = predict(email)
print(f" {email!r:28s} -> 【{label}】")
EXAMPLE 朴素贝叶斯分类结果: '免费 领取 现金 奖品' -> 【spam】 '明天 会议 项目 汇报' -> 【ham】 '限时 点击 优惠 链接' -> 【spam】 '周末 大家 一起 讨论 报告' -> 【ham】
Application Scenarios in AI
| Scenario | Description |
|---|---|
| Text classification | Spam filtering, news classification, sentiment analysis - Naive Bayes is a classic baseline |
| Document topic modeling | Mathematical foundation of LDA (Latent Dirichlet Allocation) |
| Anomaly detection | Calculate the probability that a new sample belongs to the "normal distribution"; if it is too low, classify it as an anomaly. |