Text Representation Methods
Text representation is a fundamental task in natural language processing (NLP), which converts unstructured text data into numerical forms that computers can process.
This article systematically introduces commonly used text representation methods in NLP, from traditional methods to modern deep learning techniques, helping readers fully understand this core concept.
Traditional Text Representation
Bag of Words Model
The bag of words model is one of the simplest text representation methods; it treats text as an unordered collection of words.
Basic Concepts
- Ignores word order and grammar, only focuses on whether words appear
- Build a vocabulary and count the occurrences of each word in the document
- The final representation is a high-dimensional sparse vector
Code Example
Example
corpus = [
'This is the first document.',
'This document is the second document.',
'And this is the third one.',
'Is this the first document?'
]
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
print(vectorizer.get_feature_names_out())
print(X.toarray())
Pros and Cons Analysis
✅ Advantages:
- Simple implementation, high computational efficiency
- Suitable for small-scale datasets and simple tasks
❌ Disadvantages:
- Ignores word order and semantic information
- High-dimensional sparsity problem
- Cannot handle synonyms and polysemous words
TF-IDF
TF-IDF (Term Frequency-Inverse Document Frequency) is an improvement over the bag of words model that considers the importance of words in the entire corpus.
Calculation Formula
- TF (Term Frequency):
The number of times a word appears in a document / total number of words in the document - IDF (Inverse Document Frequency):
log(文档总数 / 包含该词的文档数) - TF-IDF = TF × IDF
Code Implementation
Example
tfidf_vectorizer = TfidfVectorizer()
X_tfidf = tfidf_vectorizer.fit_transform(corpus)
print(tfidf_vectorizer.get_feature_names_out())
print(X_tfidf.toarray())
Pros and Cons
✅ Advantages:
- Reduces the impact of common words and highlights important words
- Better performance than the simple bag of words model
❌ Disadvantages:
- Still cannot capture semantic relationships
- The high-dimensional problem still exists
N-gram Model
The N-gram model takes word order information into account and represents text through combinations of n consecutive words.
Common Types
- Unigram (1-gram): a single word
- Bigram (2-gram): combination of two consecutive words
- Trigram (3-gram): combination of three consecutive words
Code Example
Example
X_bigram = bigram_vectorizer.fit_transform(corpus)
print(bigram_vectorizer.get_feature_names_out())
Pros and Cons
✅ Advantages:
- Captures local word order information
- Can represent phrases and fixed collocations
❌ Disadvantages:
- The dimensionality explosion problem is more severe
- Still cannot handle long-distance dependencies
Word Vector Representation
Word2Vec Principles and Implementation
Word2Vec is a neural network-based word vector representation method proposed by Google in 2013.
Two model architectures
- CBOW(Continuous Bag of Words): predict the current word from context
- Skip-gram: predict the context from the current word
Code Implementation
Example
sentences = [["cat", "say", "meow"], ["dog", "say", "woof"]]
model = Word2Vec(sentences, vector_size=100, window=5, min_count=1, workers=4)
# Get word vectors
vector = model.wv['cat']
# Find similar words
similar_words = model.wv.most_similar('cat')
Features
- Low-dimensional dense vectors (usually 50-300 dimensions)
- Can capture semantic and syntactic relationships between words
- Supports vector operations (e.g., king - man + woman ≈ queen)
GloVe Word Vectors
GloVe (Global Vectors for Word Representation) combines the advantages of global statistical information and local context windows.
Core Idea
- Based on a word co-occurrence matrix
- The optimization goal is to make the dot product of two word vectors equal to the logarithm of their co-occurrence count
Comparison with Word2Vec
| Feature | Word2Vec | GloVe |
|---|---|---|
| Training Method | Local window | Global statistics |
| Computational Efficiency | Relatively high | Relatively low |
| Performance on small datasets | Better | Average |
| Performance on large datasets | OK | Better |
FastText
FastText is a word vector model developed by Facebook, characterized by considering subword information.
Main Features
- Represents words as a set of character n-grams
- Can handle out-of-vocabulary (OOV) words
- Particularly suitable for morphologically rich languages
Code Example
Example
model = FastText(sentences, vector_size=100, window=5, min_count=1, workers=4)
# Even words not in the dictionary can get vectors
vector = model.wv['unseenword']
Context-Aware Representations
ELMo Model
ELMo (Embeddings from Language Models) is one of the earliest context-dependent word representation methods.
Core Features
- Based on bidirectional LSTM language models
- The representation of a word depends on the entire input sentence
- Generates multi-layer representations (can combine semantics at different levels)
Architecture Diagram

BERT and Its Variants
BERT (Bidirectional Encoder Representations from Transformers) is a pretrained language model proposed by Google.
Key Innovations
- Transformer architecture
- Masked Language Model (MLM) training objective
- Next Sentence Prediction (NSP) task
Common Variants
- RoBERTa: optimized training strategy
- DistilBERT: lightweight version of BERT
- ALBERT: parameter sharing reduces model size
Code Example
Example
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')
inputs = tokenizer("Hello, my dog is cute", return_tensors="pt")
outputs = model(**inputs)
last_hidden_states = outputs.last_hidden_state
Overview of Pretrained Language Models
Modern NLP mainly uses the pretraining + fine-tuning paradigm:
- Pretraining stage: train general-purpose language representations on large-scale corpora
- Fine-tuning stage: adjust model parameters on task-specific data
Model Comparison
| Model | Release Time | Main Features |
|---|---|---|
| Word2Vec | 2013 | Static word vectors |
| GloVe | 2014 | Global statistics + local window |
| ELMo | 2018 | Bidirectional LSTM, context-dependent |
| BERT | 2018 | Transformer, bidirectional context |
| GPT-3 | 2020 | Unidirectional Transformer, strong generation ability |
Document-Level Representations
Doc2Vec
Doc2Vec is an extension of Word2Vec that can directly learn vector representations of documents.
Two Models
- PV-DM(Distributed Memory): similar to CBOW, adds document ID
- PV-DBOW(Distributed Bag of Words): similar to Skip-gram
Code Example
Example
from gensim.models.doc2vec import TaggedDocument
documents = [TaggedDocument(doc, [i]) for i, doc in enumerate(corpus)]
model = Doc2Vec(documents, vector_size=100, window=5, min_count=1, workers=4)
vector = model.infer_vector(["new", "document", "text"])
Sentence Vectors and Document Vectors
Common Methods
- Averaging: take the average of word vectors
- SIF: smoothed inverse frequency weighted average
- BERT sentence vectors: use the [CLS] token or average all word vectors
Code Example (using Sentence-BERT)
Example
model = SentenceTransformer('all-MiniLM-L6-v2')
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
Topic Model (LDA)
Latent Dirichlet Allocation (LDA) is an unsupervised topic modeling method.
Basic Principles
- Represents a document as a mixture of multiple topics
- Each topic is a probability distribution over words
- Learned via variational inference or Gibbs sampling
Code Example
Example
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
lda = LatentDirichletAllocation(n_components=2)
lda.fit(X)
Application Scenarios
- Document clustering
- Content recommendation
- Text summarization
Summary
The development of text representation methods has evolved from simple statistics to deep learning:
- Traditional methods: Simple and efficient, suitable for small-scale data
- Word vectors: Captures semantic relationships, low dimensionality
- Context-aware models: Dynamic representation, best performance but high computational cost
- Document representation: Extends from word level to document level
When choosing text representation methods, consider:
- Task requirements (whether semantic understanding is needed)
- Data scale
- Computational resources
- Language characteristics
With the development of large language models, text representation techniques are still evolving rapidly, but understanding these fundamental methods remains crucial for mastering NLP.
Other extensions