Text Preprocessing
Text preprocessing is a fundamental and critical step in Natural Language Processing (NLP), as it transforms raw unstructured text data into a format suitable for machine learning model processing.
This article systematically introduces the three core steps of text preprocessing: text cleaning, tokenization, and part-of-speech tagging.
Text Cleaning: Purifying Raw Text Data
Text cleaning is the first step in preprocessing, aiming to remove noise data from the text and improve the accuracy of subsequent processing.
Encoding Format Handling
Text from different sources may use different encoding formats (such as UTF-8, GBK, ASCII, etc.), so unifying the encoding is the top priority:
Example
text = "example text".encode('gbk') # Assume the original encoding is GBK
text = text.decode('gbk').encode('utf-8') # Convert to UTF-8
Common solutions to encoding issues:
- Use
chardetlibrary to automatically detect encoding - Uniformly convert to UTF-8 encoding
- Handle undecodable characters (usually replace or ignore them)
Special Character Handling
Different types of special characters need to be handled in different scenarios:
| Character Type | Handling Method | Application Scenario |
|---|---|---|
| HTML tags | Remove with regular expressions | Web-scraped text |
| Emojis | Remove or convert to text descriptions | Social media analysis |
| Control characters | Filter out | All text processing |
| Special punctuation | Standardize | Text normalization |
Example
# Example of removing HTML tags
text = "<p>This is an <b>HTML</b> text</p>"
clean_text = re.sub(r'<[^>]+>', '', text)
print(clean_text) # Output: This is an HTML text
Noise Data Removal
Depending on the specific task requirements, it may be necessary to:
- Remove irrelevant information (ads, copyright notices, etc.)
- Handle spelling errors (use spell-checking libraries)
- Standardize numeric representations (e.g., unify "1000" as "1,000")
- Unify date formats ("2023-01-01" vs. "01/01/2023")
Tokenization: Breaking Text into Basic Units
Tokenization is the process of splitting continuous text into meaningful linguistic units (tokens). Different languages require different tokenization methods.
English Tokenization Methods
English tokenization is relatively simple, mainly based on splitting by spaces and punctuation:
Example
from nltk.tokenize import word_tokenize
text = "Natural Language Processing is fascinating!"
tokens = word_tokenize(text)
print(tokens) # ['Natural', 'Language', 'Processing', 'is', 'fascinating', '!']
Notes on English tokenization:
- Handle contractions (e.g., "I'm" → "I" + "'m")
- Preserve or merge specific phrases (e.g., "New York" as one token)
- Handle hyphens ("state-of-the-art")
Chinese Word Segmentation Techniques
Chinese has no obvious word boundaries, making word segmentation more complex. The main methods include:
- Dictionary-based word segmentation: maximum matching, shortest path methods
- Statistical word segmentation: sequence labeling methods such as HMM and CRF
- Deep learning-based word segmentation: models such as BiLSTM-CRF and BERT
Example
import jieba
text = Natural language processing is very interesting.
tokens = jieba.lcut(text)
print(tokens) # ['natural language', 'processing', 'very', 'interesting']
Subword Tokenization
To solve the problems of rare words and vocabulary explosion, common methods include:
- **Byte Pair Encoding (BPE)**: builds subwords by merging high-frequency character pairs
- WordPiece: similar to BPE, but merges based on probability
- Unigram Language Model: starts from a large vocabulary and progressively removes low-probability subwords
Example
from transformers import BertTokenizer
tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
tokens = tokenizer.tokenize("Natural language processing")
print(tokens) # ['自', '然', '语', '言', '处', '理']
Comparison of Common Tokenization Tools
| Tool Name | Supported Languages | Features | Applicable Scenarios |
|---|---|---|---|
| NLTK | Primarily English | Comprehensive features, moderate speed | Teaching, research |
| spaCy | Multilingual | Industrial-grade, fast | Production environment |
| jieba | Chinese | Simple and easy to use, extensible dictionaries | Chinese text processing |
| Stanford CoreNLP | Multilingual | High accuracy, resource-intensive | Academic research |
| HuggingFace Tokenizers | Multilingual | Supports subword tokenization | Deep learning |
Part-of-Speech Tagging: Understanding the Grammatical Role of Words
Part-of-Speech Tagging is the process of assigning a part-of-speech category to each word in the tokenization result.
Concept of Part-of-Speech Tagging
Part-of-speech tagging helps to:
- Understand sentence structure
- Resolve word sense ambiguity
- Support more advanced NLP tasks (such as syntactic parsing)
Common Part-of-Speech Tag Sets
Different languages and tools use different part-of-speech tagging systems:
Commonly used Penn Treebank tag set for English (partial):
- NN: noun
- VB: verb
- JJ: adjective
- RB: adverb
- PRP: pronoun
Commonly used ICTCLAS tag set for Chinese (partial):
- n: noun
- v: verb
- a: adjective
- d: adverb
- r: pronoun
Automatic Part-of-Speech Tagging Methods
- Rule-based methods: tagging using manually written rules
- Statistical methods: models such as HMM and MaxEnt
- Deep learning-based methods: neural networks such as RNN and Transformer
Example
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Natural Language Processing is fascinating!")
for token in doc:
print(token.text, token.pos_) # Output each word and its POS tag
Evaluation metrics for part-of-speech tagging:
- Accuracy
- Unknown word accuracy (OOV Accuracy)
- Confusion matrix analysis
Practical Recommendations
- Preprocessing pipeline order: encoding handling → text cleaning → tokenization → part-of-speech tagging
- Tool selection principles: choose appropriate tools based on language, task requirements, and performance needs
- Custom processing: custom dictionaries or rules may be required for specific domains
- Performance optimization: for large-scale text, consider using parallel processing or efficient tools