Naive Bayes Spam Classifier

Implement a simple Chinese spam classifier using the most basic conditional probability formula, without calling sklearn.

After completing this case, you will understand:Bayes' theorem + the "word independence" assumption = a usable text classifier.


Everyday Introduction

Email spam filtering

Your mailbox receives 100 emails every day. How can the program automatically determine whether a new email is spam? The program's approach: count, in historical emails, the probability of words like "free" and "winning" appearing in spam, and the probability of words like "meeting" and "project" appearing in normal emails.

When a new email arrives, calculate: the probability of these words appearing together under the spam assumption vs. under the normal email assumption. Whichever is larger wins the classification.


Intuitive Understanding

The key simplification of Naive Bayes — the "naive assumption": it assumes that the occurrence of each word is independent of the others. Although this assumption does not hold in reality ("free" and "prize" often appear together), its practical performance is surprisingly good.

TechniquesApproachProblem Solved
Logarithmic ProbabilityUse log to convert multiplication into additionMultiplying many small decimals can become 0 (numerical underflow)
Laplace SmoothingAdd 1 to each word count instead of 0Words not seen in the training set cause probability to be 0

Python Hands-On Practice

Example

from collections import defaultdict
import math

# Training data (8 labeled Chinese SMS messages)
train_data = [
    (Win Free Prizes Click Now, "spam"),
    (Congratulations Winning Claim Cash Discount, "spam"),
    (Limited Time Offer Click Link Free, "spam"),
    (Loan Low Interest Apply Now Cash, "spam"),
    (Meeting Arrangement Tomorrow Morning Discussion, "ham"),
    (Project Progress Report Meeting Minutes, "ham"),
    (Weekend Dinner Arrangement Everyone Attend, "ham"),
    (Report Data Analysis Project Completion, "ham"),
]

# Count word frequencies for each category
class_word_counts = defaultdict(lambda: defaultdict(int))
class_doc_counts = defaultdict(int)
vocab = set()

for text, label in train_data:
    class_doc_counts[label] += 1
    for word in text.split():
        class_word_counts[label][word] += 1
        vocab.add(word)

vocab_size = len(vocab)
total_docs = len(train_data)

# Naive Bayes classification (log probability + Laplace smoothing)
def predict(text, alpha=1.0):
    words = text.split()
    scores = {}
    for label in class_doc_counts:
        log_prob = math.log(class_doc_counts[label] / total_docs)
        total_words = sum(class_word_counts[label].values())
        for word in words:
            word_count = class_word_counts[label][word]
            word_prob = (word_count + alpha) / (total_words + alpha * vocab_size)
            log_prob += math.log(word_prob)
        scores[label] = log_prob
    return max(scores, key=scores.get), scores

# Test
test_emails = [
    Claim Free Cash Prizes,
    Tomorrow's meeting project report,
    Time-limited discount link,
    On the weekend, everyone discusses the report together.,
]

print(EXAMPLE Naive Bayes classification result:\n")
for email in test_emails:
    label, scores = predict(email)
    print(f"  {email!r:28s} -> 【{label}】")
EXAMPLE 朴素贝叶斯分类结果:

  '免费 领取 现金 奖品'          -> 【spam】
  '明天 会议 项目 汇报'           -> 【ham】
  '限时 点击 优惠 链接'           -> 【spam】
  '周末 大家 一起 讨论 报告'       -> 【ham】

Application Scenarios in AI

ScenarioDescription
Text classificationSpam filtering, news classification, sentiment analysis - Naive Bayes is a classic baseline
Document topic modelingMathematical foundation of LDA (Latent Dirichlet Allocation)
Anomaly detectionCalculate the probability that a new sample belongs to the "normal distribution"; if it is too low, classify it as an anomaly.
Other extensions