Text Representation Methods

Text representation is a fundamental task in natural language processing (NLP), which converts unstructured text data into numerical forms that computers can process.

This article systematically introduces commonly used text representation methods in NLP, from traditional methods to modern deep learning techniques, helping readers fully understand this core concept.


Traditional Text Representation

Bag of Words Model

The bag of words model is one of the simplest text representation methods; it treats text as an unordered collection of words.

Basic Concepts

  • Ignores word order and grammar, only focuses on whether words appear
  • Build a vocabulary and count the occurrences of each word in the document
  • The final representation is a high-dimensional sparse vector

Code Example

Example

from sklearn.feature_extraction.text import CountVectorizer

corpus = [
    'This is the first document.',
    'This document is the second document.',
    'And this is the third one.',
    'Is this the first document?'
]

vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
print(vectorizer.get_feature_names_out())
print(X.toarray())

Pros and Cons Analysis

✅ Advantages:

  • Simple implementation, high computational efficiency
  • Suitable for small-scale datasets and simple tasks

❌ Disadvantages:

  • Ignores word order and semantic information
  • High-dimensional sparsity problem
  • Cannot handle synonyms and polysemous words

TF-IDF

TF-IDF (Term Frequency-Inverse Document Frequency) is an improvement over the bag of words model that considers the importance of words in the entire corpus.

Calculation Formula

  • TF (Term Frequency):The number of times a word appears in a document / total number of words in the document
  • IDF (Inverse Document Frequency):log(文档总数 / 包含该词的文档数)
  • TF-IDF = TF × IDF

Code Implementation

Example

from sklearn.feature_extraction.text import TfidfVectorizer

tfidf_vectorizer = TfidfVectorizer()
X_tfidf = tfidf_vectorizer.fit_transform(corpus)
print(tfidf_vectorizer.get_feature_names_out())
print(X_tfidf.toarray())

Pros and Cons

✅ Advantages:

  • Reduces the impact of common words and highlights important words
  • Better performance than the simple bag of words model

❌ Disadvantages:

  • Still cannot capture semantic relationships
  • The high-dimensional problem still exists

N-gram Model

The N-gram model takes word order information into account and represents text through combinations of n consecutive words.

Common Types

  • Unigram (1-gram): a single word
  • Bigram (2-gram): combination of two consecutive words
  • Trigram (3-gram): combination of three consecutive words

Code Example

Example

bigram_vectorizer = CountVectorizer(ngram_range=(2, 2))
X_bigram = bigram_vectorizer.fit_transform(corpus)
print(bigram_vectorizer.get_feature_names_out())

Pros and Cons

✅ Advantages:

  • Captures local word order information
  • Can represent phrases and fixed collocations

❌ Disadvantages:

  • The dimensionality explosion problem is more severe
  • Still cannot handle long-distance dependencies

Word Vector Representation

Word2Vec Principles and Implementation

Word2Vec is a neural network-based word vector representation method proposed by Google in 2013.

Two model architectures

  1. CBOW(Continuous Bag of Words): predict the current word from context
  2. Skip-gram: predict the context from the current word

Code Implementation

Example

from gensim.models import Word2Vec

sentences = [["cat", "say", "meow"], ["dog", "say", "woof"]]
model = Word2Vec(sentences, vector_size=100, window=5, min_count=1, workers=4)

# Get word vectors
vector = model.wv['cat']
# Find similar words
similar_words = model.wv.most_similar('cat')

Features

  • Low-dimensional dense vectors (usually 50-300 dimensions)
  • Can capture semantic and syntactic relationships between words
  • Supports vector operations (e.g., king - man + woman ≈ queen)

GloVe Word Vectors

GloVe (Global Vectors for Word Representation) combines the advantages of global statistical information and local context windows.

Core Idea

  • Based on a word co-occurrence matrix
  • The optimization goal is to make the dot product of two word vectors equal to the logarithm of their co-occurrence count

Comparison with Word2Vec

Feature Word2Vec GloVe
Training Method Local window Global statistics
Computational Efficiency Relatively high Relatively low
Performance on small datasets Better Average
Performance on large datasets OK Better

FastText

FastText is a word vector model developed by Facebook, characterized by considering subword information.

Main Features

  • Represents words as a set of character n-grams
  • Can handle out-of-vocabulary (OOV) words
  • Particularly suitable for morphologically rich languages

Code Example

Example

from gensim.models import FastText

model = FastText(sentences, vector_size=100, window=5, min_count=1, workers=4)
# Even words not in the dictionary can get vectors
vector = model.wv['unseenword']

Context-Aware Representations

ELMo Model

ELMo (Embeddings from Language Models) is one of the earliest context-dependent word representation methods.

Core Features

  • Based on bidirectional LSTM language models
  • The representation of a word depends on the entire input sentence
  • Generates multi-layer representations (can combine semantics at different levels)

Architecture Diagram


BERT and Its Variants

BERT (Bidirectional Encoder Representations from Transformers) is a pretrained language model proposed by Google.

Key Innovations

  • Transformer architecture
  • Masked Language Model (MLM) training objective
  • Next Sentence Prediction (NSP) task

Common Variants

  1. RoBERTa: optimized training strategy
  2. DistilBERT: lightweight version of BERT
  3. ALBERT: parameter sharing reduces model size

Code Example

Example

from transformers import BertTokenizer, BertModel

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')
model = BertModel.from_pretrained('bert-base-uncased')

inputs = tokenizer("Hello, my dog is cute", return_tensors="pt")
outputs = model(**inputs)
last_hidden_states = outputs.last_hidden_state

Overview of Pretrained Language Models

Modern NLP mainly uses the pretraining + fine-tuning paradigm:

  1. Pretraining stage: train general-purpose language representations on large-scale corpora
  2. Fine-tuning stage: adjust model parameters on task-specific data

Model Comparison

Model Release Time Main Features
Word2Vec 2013 Static word vectors
GloVe 2014 Global statistics + local window
ELMo 2018 Bidirectional LSTM, context-dependent
BERT 2018 Transformer, bidirectional context
GPT-3 2020 Unidirectional Transformer, strong generation ability

Document-Level Representations

Doc2Vec

Doc2Vec is an extension of Word2Vec that can directly learn vector representations of documents.

Two Models

  1. PV-DM(Distributed Memory): similar to CBOW, adds document ID
  2. PV-DBOW(Distributed Bag of Words): similar to Skip-gram

Code Example

Example

from gensim.models import Doc2Vec
from gensim.models.doc2vec import TaggedDocument

documents = [TaggedDocument(doc, [i]) for i, doc in enumerate(corpus)]
model = Doc2Vec(documents, vector_size=100, window=5, min_count=1, workers=4)
vector = model.infer_vector(["new", "document", "text"])

Sentence Vectors and Document Vectors

Common Methods

  1. Averaging: take the average of word vectors
  2. SIF: smoothed inverse frequency weighted average
  3. BERT sentence vectors: use the [CLS] token or average all word vectors

Code Example (using Sentence-BERT)

Example

from sentence_transformers import SentenceTransformer

model = SentenceTransformer('all-MiniLM-L6-v2')
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)

Topic Model (LDA)

Latent Dirichlet Allocation (LDA) is an unsupervised topic modeling method.

Basic Principles

  • Represents a document as a mixture of multiple topics
  • Each topic is a probability distribution over words
  • Learned via variational inference or Gibbs sampling

Code Example

Example

from sklearn.decomposition import LatentDirichletAllocation
from sklearn.feature_extraction.text import CountVectorizer

vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
lda = LatentDirichletAllocation(n_components=2)
lda.fit(X)

Application Scenarios

  • Document clustering
  • Content recommendation
  • Text summarization

Summary

The development of text representation methods has evolved from simple statistics to deep learning:

  1. Traditional methods: Simple and efficient, suitable for small-scale data
  2. Word vectors: Captures semantic relationships, low dimensionality
  3. Context-aware models: Dynamic representation, best performance but high computational cost
  4. Document representation: Extends from word level to document level

When choosing text representation methods, consider:

  • Task requirements (whether semantic understanding is needed)
  • Data scale
  • Computational resources
  • Language characteristics

With the development of large language models, text representation techniques are still evolving rapidly, but understanding these fundamental methods remains crucial for mastering NLP.

Other extensions