Text Classification
Text Classification is one of the most fundamental and important tasks in Natural Language Processing (NLP). Its goal is to automatically categorize a given text document into one or more predefined categories.
Basic Concepts
Text classification is like a librarian who needs to categorize books and place them on the correct shelves based on their content. In the field of computing, we need to teach machines how to understand text content and make correct classification decisions.
Application Scenarios
Text classification has a wide range of applications in modern society:
- Sentiment Analysis: Determine whether a review is positive or negative
- Spam Filtering: Distinguish between normal emails and spam
- News Classification: Categorize news into sections such as sports, finance, and technology
- Intent Recognition: Understand the true intent of user queries
- Medical Diagnosis: Classify disease types based on symptom descriptions
Basic Workflow of Text Classification
A complete text classification system typically includes the following steps:

1. Text Preprocessing
Text preprocessing transforms raw text into a form suitable for machine learning models to process:
Example
import nltk
from nltk.corpus import stopwords
from nltk.stem import PorterStemmer
def preprocess_text(text):
# Convert to lowercase
text = text.lower()
# Remove special characters and numbers
text = re.sub(r'[^a-zA-Z\s]', '', text)
# Tokenization
words = text.split()
# Remove stop words
stop_words = set(stopwords.words('english'))
words = [word for word in words if word not in stop_words]
# Stemming
stemmer = PorterStemmer()
words = [stemmer.stem(word) for word in words]
return ' '.join(words)
2. Feature Extraction
Convert text into numerical feature representations. Common methods include:
| Method | Description | Advantages | Disadvantages |
|---|---|---|---|
| Bag of Words (BoW) | Counts word frequency | Simple and intuitive | Ignores word order and semantics |
| TF-IDF | Considers word importance | More accurate than BoW | Still ignores context |
| Word2Vec | Word Vector Representations | Captures semantic relationships | Cannot handle polysemous words |
| BERT | Contextual Embeddings | State-of-the-art representation | High computational resource requirements |
3. Classification Model Selection
Choose an appropriate classification algorithm based on task requirements and data characteristics:
Traditional Machine Learning Methods:
- Naive Bayes
- Support Vector Machine (SVM)
- Logistic Regression
- Random Forest
Deep Learning Methods:
- Convolutional Neural Network (CNN)
- Recurrent Neural Network (RNN/LSTM)
- Transformer Models (BERT, etc.)
Practical Example: News Classification
Let's demonstrate how to implement text classification using Python with a practical example. We will use the 20 Newsgroups dataset, a classic news classification dataset.
1. Data Preparation
Example
# Select 4 categories as an example
categories = ['alt.atheism', 'soc.religion.christian', 'comp.graphics', 'sci.med']
# Load training and test sets
newsgroups_train = fetch_20newsgroups(subset='train', categories=categories)
newsgroups_test = fetch_20newsgroups(subset='test', categories=categories)
print(f"Training set sample count: {len(newsgroups_train.data)}")
print(f"Test set sample count: {len(newsgroups_test.data)}")
2. Feature Extraction (TF-IDF)
Example
# Create a TF-IDF vectorizer
vectorizer = TfidfVectorizer(max_features=5000)
# Transform the training and test sets
X_train = vectorizer.fit_transform(newsgroups_train.data)
X_test = vectorizer.transform(newsgroups_test.data)
y_train = newsgroups_train.target
y_test = newsgroups_test.target
3. Model Training (Logistic Regression)
Example
from sklearn.metrics import accuracy_score, classification_report
# Create and train the model
model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
# Predict on the test set
y_pred = model.predict(X_test)
# Evaluate the model
print(f"Accuracy: {accuracy_score(y_test, y_pred):.2f}")
print("\nClassification report:)
print(classification_report(y_test, y_pred, target_names=newsgroups_test.target_names))
4. Result Analysis
A typical output might look like this:
准确率: 0.91
分类报告:
precision recall f1-score support
alt.atheism 0.90 0.87 0.89 319
soc.religion.christian 0.93 0.95 0.94 389
comp.graphics 0.89 0.90 0.90 396
sci.med 0.92 0.91 0.92 398
accuracy 0.91 1502
macro avg 0.91 0.91 0.91 1502
weighted avg 0.91 0.91 0.91 1502
Advanced Techniques and Challenges
Handling Class Imbalance
When some categories have far more samples than others, you can try:
- Resampling (oversampling the minority class or undersampling the majority class)
- Using class weights
- Trying different evaluation metrics (such as F1-score instead of accuracy)
Methods to Improve Model Performance
Feature Engineering:
- Try different n-gram ranges
- Add part-of-speech features
- Use more advanced word embeddings
Model Optimization:
- Hyperparameter tuning
- Model ensembling
- Try deep learning models
Data Augmentation:
- Back Translation
- Synonym replacement
- Generative Adversarial Network (GAN)
Common Challenges
- Multi-label classification: A document may belong to multiple categories
- Domain adaptation: Model performance degrades in new domains
- Few-shot learning: Situations with limited labeled data
- Interpretability: Understanding why the model makes a particular classification decision
Summary and Learning Path
Text classification is a fundamental NLP task, and mastering it is crucial for understanding more complex NLP applications. Recommended learning path:
- Start with traditional machine learning methods (such as TF-IDF + SVM)
- Understand the concepts and applications of word embeddings (Word2Vec, GloVe)
- Learn about the application of deep learning models (CNN, LSTM) in text classification
- Explore transfer learning with pre-trained language models (BERT, GPT)
Through continuous practice and trying different methods, you will be able to build powerful and practical text classification systems.
Other Extensions