Text Preprocessing

Text preprocessing is a fundamental and critical step in Natural Language Processing (NLP), as it transforms raw unstructured text data into a format suitable for machine learning model processing.

This article systematically introduces the three core steps of text preprocessing: text cleaning, tokenization, and part-of-speech tagging.


Text Cleaning: Purifying Raw Text Data

Text cleaning is the first step in preprocessing, aiming to remove noise data from the text and improve the accuracy of subsequent processing.

Encoding Format Handling

Text from different sources may use different encoding formats (such as UTF-8, GBK, ASCII, etc.), so unifying the encoding is the top priority:

Example

# Encoding conversion example
text = "example text".encode('gbk')  # Assume the original encoding is GBK
text = text.decode('gbk').encode('utf-8')  # Convert to UTF-8

Common solutions to encoding issues:

  • Usechardetlibrary to automatically detect encoding
  • Uniformly convert to UTF-8 encoding
  • Handle undecodable characters (usually replace or ignore them)

Special Character Handling

Different types of special characters need to be handled in different scenarios:

Character Type Handling Method Application Scenario
HTML tags Remove with regular expressions Web-scraped text
Emojis Remove or convert to text descriptions Social media analysis
Control characters Filter out All text processing
Special punctuation Standardize Text normalization

Example

import re

# Example of removing HTML tags
text = "<p>This is an <b>HTML</b> text</p>"
clean_text = re.sub(r'<[^>]+>', '', text)
print(clean_text)  # Output: This is an HTML text

Noise Data Removal

Depending on the specific task requirements, it may be necessary to:

  1. Remove irrelevant information (ads, copyright notices, etc.)
  2. Handle spelling errors (use spell-checking libraries)
  3. Standardize numeric representations (e.g., unify "1000" as "1,000")
  4. Unify date formats ("2023-01-01" vs. "01/01/2023")

Tokenization: Breaking Text into Basic Units

Tokenization is the process of splitting continuous text into meaningful linguistic units (tokens). Different languages require different tokenization methods.

English Tokenization Methods

English tokenization is relatively simple, mainly based on splitting by spaces and punctuation:

Example

# Tokenize English with NLTK
from nltk.tokenize import word_tokenize

text = "Natural Language Processing is fascinating!"
tokens = word_tokenize(text)
print(tokens)  # ['Natural', 'Language', 'Processing', 'is', 'fascinating', '!']

Notes on English tokenization:

  • Handle contractions (e.g., "I'm" → "I" + "'m")
  • Preserve or merge specific phrases (e.g., "New York" as one token)
  • Handle hyphens ("state-of-the-art")

Chinese Word Segmentation Techniques

Chinese has no obvious word boundaries, making word segmentation more complex. The main methods include:

  1. Dictionary-based word segmentation: maximum matching, shortest path methods
  2. Statistical word segmentation: sequence labeling methods such as HMM and CRF
  3. Deep learning-based word segmentation: models such as BiLSTM-CRF and BERT

Example

# Use jieba for Chinese word segmentation
import jieba

text = Natural language processing is very interesting.
tokens = jieba.lcut(text)
print(tokens)  # ['natural language', 'processing', 'very', 'interesting']

Subword Tokenization

To solve the problems of rare words and vocabulary explosion, common methods include:

  • **Byte Pair Encoding (BPE)**: builds subwords by merging high-frequency character pairs
  • WordPiece: similar to BPE, but merges based on probability
  • Unigram Language Model: starts from a large vocabulary and progressively removes low-probability subwords

Example

# Example of using HuggingFace's tokenizer
from transformers import BertTokenizer

tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
tokens = tokenizer.tokenize("Natural language processing")
print(tokens)  # ['自', '然', '语', '言', '处', '理']

Comparison of Common Tokenization Tools

Tool Name Supported Languages Features Applicable Scenarios
NLTK Primarily English Comprehensive features, moderate speed Teaching, research
spaCy Multilingual Industrial-grade, fast Production environment
jieba Chinese Simple and easy to use, extensible dictionaries Chinese text processing
Stanford CoreNLP Multilingual High accuracy, resource-intensive Academic research
HuggingFace Tokenizers Multilingual Supports subword tokenization Deep learning

Part-of-Speech Tagging: Understanding the Grammatical Role of Words

Part-of-Speech Tagging is the process of assigning a part-of-speech category to each word in the tokenization result.

Concept of Part-of-Speech Tagging

Part-of-speech tagging helps to:

  • Understand sentence structure
  • Resolve word sense ambiguity
  • Support more advanced NLP tasks (such as syntactic parsing)

Common Part-of-Speech Tag Sets

Different languages and tools use different part-of-speech tagging systems:

Commonly used Penn Treebank tag set for English (partial):

  • NN: noun
  • VB: verb
  • JJ: adjective
  • RB: adverb
  • PRP: pronoun

Commonly used ICTCLAS tag set for Chinese (partial):

  • n: noun
  • v: verb
  • a: adjective
  • d: adverb
  • r: pronoun

Automatic Part-of-Speech Tagging Methods

  1. Rule-based methods: tagging using manually written rules
  2. Statistical methods: models such as HMM and MaxEnt
  3. Deep learning-based methods: neural networks such as RNN and Transformer

Example

# Use spaCy for part-of-speech tagging
import spacy

nlp = spacy.load("en_core_web_sm")
doc = nlp("Natural Language Processing is fascinating!")
for token in doc:
    print(token.text, token.pos_)  # Output each word and its POS tag

Evaluation metrics for part-of-speech tagging:

  • Accuracy
  • Unknown word accuracy (OOV Accuracy)
  • Confusion matrix analysis

Practical Recommendations

  1. Preprocessing pipeline order: encoding handling → text cleaning → tokenization → part-of-speech tagging
  2. Tool selection principles: choose appropriate tools based on language, task requirements, and performance needs
  3. Custom processing: custom dictionaries or rules may be required for specific domains
  4. Performance optimization: for large-scale text, consider using parallel processing or efficient tools
Other Extensions