BERT Series Models
BERT (Bidirectional Encoder Representations from Transformers) is a revolutionary natural language processing model proposed by Google in 2018, which completely changed the research and application paradigms in the NLP field.
This article will systematically introduce BERT's core principles, training methods, fine-tuning techniques, and mainstream variant models.
BERT Architecture and Training
The figure below showsBERT(Bidirectional Encoder Representations from Transformers)the model's core architecture and the masked language modeling (Masked Language Modeling, MLM) task during pretraining.

1. Input Layer (Embedding)
- Input sequence: text composed of words (or subwords), for example
[W₁, W₂, W₃, [MASK], W₅, W₆, W₇, W₂, W₃, W₄, W₅]。[MASK]is the word randomly masked by BERT during pretraining (as in the original text,W₄is replaced by[MASK])。
- Embedding layer: converts each word into a fixed-dimensional vector representation (e.g., 768 dimensions), including:
- Token embeddings: semantic information of the vocabulary.
- Position embeddings: position information of words in the sequence.
- Segment embeddings: distinguish sentences (useful for sentence-pair tasks, not explicitly shown in the figure).
2. Transformer Encoder
- Multi-layer Transformer blocks: details not expanded in the figure, but each block contains:
- Self-attention mechanism: bidirectionally captures contextual dependencies (BERT's core feature).
- Feed-forward network: nonlinear transformation.
- Residual connections and layer normalization: stabilize the training process.
- Output: the context-aware vector representation corresponding to each input word (e.g.,
O₁, O₂, ..., O₅)。
3. Masked Language Modeling (MLM) Task
- Objective: predict the masked word
[MASK]corresponding to the original word (in the figureW₄)。 - Classification layer:
- Fully-connected layer: maps the Transformer output vector (e.g.,
O₄) to the dimension of the vocabulary size. - Activation function GELU: Gaussian Error Linear Unit (the nonlinear function used by BERT).
- Layer normalization (Norm): standardizes the output.
- Softmax: computes the probability of each word in the vocabulary and selects the word with the highest probability as the prediction result (e.g.,
W'₁, W'₂, ..., W'₅is a candidate word).
- Fully-connected layer: maps the Transformer output vector (e.g.,
Transformer Encoder Structure
BERT is built upon the encoder part of the Transformer, whose core is the multi-layer self-attention mechanism:
Example
class TransformerEncoderLayer(nn.Module):
def __init__(self, d_model, nhead, dim_feedforward=2048):
super().__init__()
self.self_attn = MultiheadAttention(d_model, nhead)
self.linear1 = nn.Linear(d_model, dim_feedforward)
self.linear2 = nn.Linear(dim_feedforward, d_model)
self.norm1 = nn.LayerNorm(d_model)
self.norm2 = nn.LayerNorm(d_model)
def forward(self, src):
# Self-attention mechanism
src2 = self.self_attn(src, src, src)[0]
src = src + self.norm1(src2)
# Feed-forward network
src2 = self.linear2(F.relu(self.linear1(src)))
src = src + self.norm2(src2)
return src
Key Innovation: Bidirectional Context Modeling
Unlike traditional language models, BERT achieves bidirectional context understanding through the following two pretraining tasks:
- Masked Language Model (MLM): randomly masks 15% of input tokens and predicts the masked words
- Next Sentence Prediction (NSP): determines whether two sentences appear consecutively
Training Parameters and Configuration
| Parameter | BERT-base | BERT-large |
|---|---|---|
| Number of layers | 12 | 24 |
| Hidden layer size | 768 | 1024 |
| Number of attention heads | 12 | 16 |
| Total parameter count | 110M | 340M |
BERT Fine-tuning Methods
Standard Fine-tuning Process
- Adding task-specific adaptation layer: add a classification/regression layer according to the downstream task
- Learning rate setting: usually use a small learning rate (2e-5 to 5e-5)
- Batch size: 16 or 32 are common choices
- Training epochs: 2-4 epochs are usually sufficient
Efficient Fine-tuning Techniques
Example
from transformers import BertForSequenceClassification, Trainer
model = BertForSequenceClassification.from_pretrained('bert-base-uncased')
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset
)
trainer.train()
Comparison of Common Fine-tuning Strategies
| Method | Advantage | Disadvantage |
|---|---|---|
| Full-parameter fine-tuning | Best performance | High computational cost |
| Feature extraction (freeze BERT) | Computationally efficient | Suboptimal performance |
| Adapter | Parameter-efficient | Requires architecture modification |
| Prompt learning | Good for few-shot performance | Requires designing prompt templates |
Mainstream BERT Variant Models
RoBERTa (Robustly Optimized BERT)
- Improvement points:
- Larger batch size (8k vs 256)
- Longer training time
- Removed NSP task
- Dynamic masking pattern
- Performance: average improvement of 2-3% on the GLUE benchmark
ALBERT (A Lite BERT)
- Core innovations:
- Parameter sharing (sharing attention parameters across layers)
- Embedding factorization (decomposing the word embedding into two small matrices)
- Effect: parameter count reduced by 89%, speed increased by 1.7 times
Other Important Variants
- DistilBERT: compresses the model through knowledge distillation
- ELECTRA: replaces MLM with a generator-discriminator architecture
- SpanBERT: optimizes modeling of text spans
Chinese BERT Models
Overview of Chinese Pretrained Models
| Model | Institution | Features |
|---|---|---|
| BERT-wwm | Harbin Institute of Technology | Whole Word Masking |
| RoBERTa-wwm-ext | Harbin Institute of Technology | Expanded training data |
| ERNIE (Baidu) | Baidu | Incorporates knowledge graphs |
| NEZHA | Huawei | Relative position encoding |
Chinese BERT Usage Example
Example
tokenizer = BertTokenizer.from_pretrained('bert-base-chinese')
model = BertModel.from_pretrained('bert-base-chinese')
inputs = tokenizer("Natural language processing is very interesting", return_tensors="pt")
outputs = model(**inputs)
Fine-tuning Recommendations for Chinese Tasks
- Using the Whole Word Masking (wwm) version yields better results
- Pay attention to handling Chinese word segmentation boundary issues
- For specialized domains, consider domain-adaptive pretraining
Practical Recommendations and Resources
Learning Roadmap

Recommended Resources
- Papers:
- Original BERT paper (arXiv:1810.04805)
- Papers on variants such as RoBERTa, ALBERT, etc.
- Code repositories:
- HuggingFace Transformers
- GitHub implementation of Chinese BERT
- Online courses:
- Coursera Natural Language Processing Specialization
- Hung-yi Lee's Deep Learning Course
Through systematic learning and practice, BERT series models can become a powerful tool for solving NLP problems. It is recommended to start with the basic version and gradually explore more advanced variants and optimization techniques.
Other Extensions