Advanced NLP Techniques

We actually interact with NLP every day—the word suggestions when typing on your phone's keyboard, the spam emails automatically filtered by your email app, and the foreign-language websites that translation software helps you understand. These are all applications of Natural Language Processing (NLP).

Before ChatGPT emerged, NLP had already developed for decades, but the technical threshold was high. You needed to understand a long list of concepts such as tokenization, part-of-speech tagging, syntactic parsing, and semantic role labeling, and also manually design features, in order to let machines "understand" a little bit of language.

Today, large language models make all of this simple. But understanding the technical evolution of NLP allows you to more thoroughly understand why large models can do these things, and where their limitations lie.

This module will take you from word vectors all the way to today's large language models, building a complete NLP knowledge system.。

Learning path: word vectors → pre-trained models → the three major paradigms of BERT/GPT/T5 → downstream tasks (classification, NER, translation, summarization). Each step has corresponding code examples to help you put theory into practice.


The Evolution of Pretrained Language Models

The development of NLP can be clearly divided into several stages, each with landmark technological breakthroughs.

Word2Vec: The Revolution of Word Embeddings

Before 2013, the way computers processed text was very primitive.

It usually used "One-Hot Encoding": each word corresponds to an extremely long vector, with only one position being 1 and the rest being 0. For example, if the vocabulary has 100,000 words, each word is a 100,000-dimensional vector.

The problem with this approach is obvious: the vectors contain no semantic information. The distance between "cat" and "dog" is the same as the distance between "cat" and "table".

The core idea of Word2Vec is:The meaning of a word is defined by the words around it.。

Example

# ============================================
# Word2Vec Basic Concept Demonstration
# Use the gensim library to train a simple word vector model
# ============================================

# First install gensim: pip install gensim

from gensim.models import Word2Vec
import numpy as np

# Prepare training data: some simple sentences
sentences = [
    ["I", "like", "eat", "apple"],
    ["I", "like", "eat", "banana"],
    ["cat", "like", "eat", "fish"],
    ["dog", "like", "eat", "meat"],
    ["apple", "is", "a kind of", "fruit"],
    ["banana", "is", "a kind of", "fruit"],
    ["cat", "is", "a kind of", "animal"],
    ["dog", "is", "a kind of", "animal"],
    ["EXAMPLE", "is", "a", Programming, Website],
    [Learning, Programming, Go, "EXAMPLE"],
]

# Train Word2Vec model
# vector_size: dimension of word vectors
# window: context window size (look at a few words before and after)
# min_count: ignore words with occurrence count less than this value
# workers: number of threads for parallel training
model = Word2Vec(
    sentences=sentences,
    vector_size=50,  # each word is represented by a 50-dimensional vector
    window=3,       # look at 3 words before and after
    min_count=1,    # keep all words
    workers=4,
    epochs=100      # train for 100 epochs
)

# get word vector
apple_vector = model.wv[apple]
print(f'apple' word vector (first 10 dimensions): {apple_vector[:10]})
print(fWord vector dimension: {len(apple_vector)})

# compute similarity between words
similarity = model.wv.similarity(apple, banana)
print(f'apple' and 'banana' similarity: {similarity:.4f})

similarity = model.wv.similarity(apple, cat)
print(f'apple' and 'cat' similarity: {similarity:.4f})

# find the most similar words
print("\nWords most similar to 'cat':)
for word, score in model.wv.most_similar(cat, topn=3):
    print(f"  {word}: {score:.4f}")

print("\nWords most similar to 'EXAMPLE':)
for word, score in model.wv.most_similar("EXAMPLE", topn=3):
    print(f"  {word}: {score:.4f}")

# classic word vector arithmetic: king - man + woman ≈ queen
# try in our small corpus: fruit - apple + fish ≈ ?
if apple in model.wv and fish in model.wv and fruit in model.wv:
    result = model.wv.most_similar(positive=[fruit, fish], negative=[apple], topn=3)
    print("\n'fruit' - 'apple' + 'fish' ≈)
    for word, score in result:
        print(f"  {word}: {score:.4f}")

The success of Word2Vec proves one thing:Semantics can be represented using vector spaces.。

But it has a limitation: each word has only one fixed vector, regardless of context. For example, "打" in "打电speech" (to make a phone call) and "打游戏" (to play games) has different meanings, but Word2Vec gives the same vector.

ELMo: Contextual Word Embeddings

ELMo (Embeddings from Language Models), introduced in 2018, solved this problem.

The idea of ELMo is:It does not pre-assign a fixed vector to each word; instead, it looks at the entire sentence and then generates a vector for that word.。

The same character “打” is one vector in “I打电speech” and another vector in “I打游戏”.

ELMo uses bidirectional LSTM (Long Short-Term Memory networks) to model context, which was the first large-scale use of the "pretraining + fine-tuning" paradigm.

GPT-1: Unidirectional Pretraining

Also in 2018, OpenAI released GPT-1 (Generative Pre-training Transformer).

Its features are:

1. Use a Transformer decoder instead of an LSTM.

2. Unidirectional: only look at the preceding words, predict the next word.

3. Generative: can continue writing text.

GPT-1 demonstrated the great potential of Transformer on NLP tasks.

BERT: Bidirectional Pretraining

At the end of 2018, Google released BERT (Bidirectional Encoder Representations from Transformers), which completely transformed the NLP field.

BERT's core innovation is:

Bidirectional: view both the preceding and following context simultaneously.

2. MLM (Masked Language Model): Randomly mask some words and let the model predict.

3. NSP (Next Sentence Prediction): Determine whether two sentences are consecutive.

BERT achieved the best results at the time on 11 NLP tasks, marking NLP's entry into the "pre-trained model era."

The Leap from GPT-3 to ChatGPT

In 2020, GPT-3 was released, with a parameter count reaching 175 billion.

People have found that when the model is large enough and the data is abundant enough, an "emergence" phenomenon occurs—the model suddenly acquires abilities that small models lack, such as few-shot learning and complex reasoning.

At the end of 2022, ChatGPT was released, using RLHF (Reinforcement Learning from Human Feedback) to make the model's outputs better align with human preferences, and AI truly went mainstream.

Let us summarize this history of evolution with a table:

yearmodelcore ideaHistorical status
2013Word2VecUse the words around a word to define its meaning.The word vector revolution: vectorized representation of semantics.
2018ELMoContext-dependent word vectorsFirst implementation of vector representation of polysemy.
2018GPT-1Transformer decoder, unidirectional pretraining.Demonstrate the potential of Transformers.
2018BERTTransformer encoder, bidirectional pre-trainingNLP enters the era of pretrained models
2020GPT-3175 billion parameters, emergent abilitiesShowcasing the infinite possibilities of large models
2022ChatGPTRLHF + conversational abilitiesAI truly reaches the masses

BERT Deep Dive

BERT is a milestone in NLP history, and it deserves our in-depth understanding.

MLM (Masked Language Model) Task

BERT's core pretraining task is MLM: randomly replace 15% of the words in a sentence with [MASK], and let the model predict what the original word was.

For example, the sentence: "I like eating apples" might become: "I [MASK] eating apples", and the model needs to predict that [MASK] is "like".

Why do this? Because it forces the model to simultaneously use information from the surrounding context, not just from the left or the right.

In practice, these 15% of the words are not all replaced with [MASK], but rather:

1. 80% of the time replace with [MASK]

2. 10% of the time replace with a random word

3. 10% of the time keep the original word unchanged

This is done to make the model more robust — it cannot be certain whether a word has been replaced, so it must truly understand the context.

NSP (Next Sentence Prediction) Task

BERT's second pretraining task is NSP: given two sentences A and B, determine whether B is the true next sentence after A.

Positive example: A="I like eating apples", B="Apples are a fruit"

Negative example: A="I like eating apples", B="The weather is really nice today"

This task helps the model understand relationships between sentences, which is very helpful for tasks such as question answering and natural language inference.

Use Cases for BERT

BERT is an "encoder" architecture, good at understanding language, suitable for:

Text classification, sentiment analysis, named entity recognition, question answering, natural language inference, etc.

It is not very good at generating text, which is GPT's strength.

Standard Process for Fine-Tuning BERT

Let's use the Hugging Face Transformers library to fine-tune BERT for text classification in practice:

Example

# ============================================
# Fine-tune BERT for text classification using Hugging Face
# Task: determine whether a sentence is positive or negative sentiment
# ============================================

# First install dependencies:
# pip install transformers datasets torch scikit-learn

import torch
from transformers import (
    BertTokenizer,
    BertForSequenceClassification,
    TrainingArguments,
    Trainer,
)
from datasets import Dataset, load_metric
import numpy as np
from sklearn.model_selection import train_test_split

# ============================================
# 1. Prepare data
# ============================================

# Construct a simple sentiment classification dataset
texts = [
    "This product is really good to use, I really like it",
    "The quality is too poor, very disappointing",
    "The EXAMPLE tutorial is clearly written and easy to learn",
    Logistics is very slow, and the packaging is damaged too,
    This phone's camera performance is excellent,
    Poor service attitude, won't come again,
    The book's content is wonderful, recommended for purchase,
    The movie is very boring, fell asleep watching it,
    This restaurant's food is very tasty,
    The game has too many bugs, poor experience,
    EXAMPLE's NLP tutorial is very helpful,
    This software's interface is well designed,
    The courier's attitude is very friendly,
    The clothing size is inaccurate, and the color is wrong too,
    The hotel environment is very quiet, slept well,
    Customer service replies very slowly, the problem wasn't resolved,
]

labels = [
    1, 0, 1, 0, 1, 0, 1, 0, 1, 0, 1, 1, 1, 0, 1, 0
]  # 1=positive, 0=negative

# Split training and validation sets
train_texts, val_texts, train_labels, val_labels = train_test_split(
    texts, labels, test_size=0.25, random_state=42
)

# Create Dataset object
train_dataset = Dataset.from_dict({"text": train_texts, "label": train_labels})
val_dataset = Dataset.from_dict({"text": val_texts, "label": val_labels})

# ============================================
# 2. Load tokenizer and model
# ============================================

# Use Chinese BERT model
model_name = "bert-base-chinese"

tokenizer = BertTokenizer.from_pretrained(model_name)
model = BertForSequenceClassification.from_pretrained(
    model_name,
    num_labels=2  # Binary classification task
)

# ============================================
# 3. Data preprocessing: convert text to model input
# ============================================

def tokenize_function(examples):
    return tokenizer(
        examples["text"],
        padding="max_length",
        truncation=True,
        max_length=64
    )

# Apply preprocessing
tokenized_train = train_dataset.map(tokenize_function, batched=True)
tokenized_val = val_dataset.map(tokenize_function, batched=True)

# ============================================
# 4. Define evaluation metrics
# ============================================

metric = load_metric("accuracy")

def compute_metrics(eval_pred):
    logits, labels = eval_pred
    predictions = np.argmax(logits, axis=-1)
    return metric.compute(predictions=predictions, references=labels)

# ============================================
# 5. Configure training parameters
# ============================================

training_args = TrainingArguments(
    output_dir="./example-bert-sentiment",  # Output directory
    num_train_epochs=3,                    # Number of training epochs
    per_device_train_batch_size=4,         # Training batch size
    per_device_eval_batch_size=4,          # Evaluation batch size
    warmup_steps=5,                        # Warmup steps
    weight_decay=0.01,                     # Weight decay
    logging_dir="./logs",                  # Log directory
    logging_steps=10,
    evaluation_strategy="epoch",           # Evaluate once per epoch
    save_strategy="epoch",
    learning_rate=2e-5,
    load_best_model_at_end=True,
)

# ============================================
# 6. Create Trainer and start training
# ============================================

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_train,
    eval_dataset=tokenized_val,
    compute_metrics=compute_metrics,
)

# Start training (commented out because this requires a long time and GPU)
# print("Start training...")
# trainer.train()

# ============================================
# 7. Use the trained model for prediction
# ============================================

# Here we directly use the original model to demonstrate the prediction process
# In actual use, you should use the model saved after training with trainer

def predict_sentiment(text, model, tokenizer):
    """Predict sentence sentiment"""
    inputs = tokenizer(
        text,
        return_tensors="pt",
        padding=True,
        truncation=True,
        max_length=64
    )

    model.eval()
    with torch.no_grad():
        outputs = model(**inputs)
        logits = outputs.logits
        probabilities = torch.softmax(logits, dim=-1)
        prediction = torch.argmax(probabilities, dim=-1).item()

    label_map = {0: "Negative", 1: "Positive"}
    confidence = probabilities[0][prediction].item()

    return {
        "text": text,
        "sentiment": label_map[prediction],
        "confidence": confidence,
        "negative_prob": probabilities[0][0].item(),
        "positive_prob": probabilities[0][1].item(),
    }

# Test a few sentences
test_texts = [
    "EXAMPLE tutorial is really great!",
    "This product's quality is too poor",
    "The weather is really nice today",
    "I don't really like this movie",
]

print("Test model prediction (using pretrained bert-base-chinese):")
print("-" * 50)
for text in test_texts:
    # Note: directly using the original model here, no fine-tuning, results are for demonstration only
    result = predict_sentiment(text, model, tokenizer)
    print(f"Text: {result['text']}")
    print(fSentiment: {result['sentiment']})
    print(fConfidence: {result['confidence']:.4f})
    print(fPositive probability: {result['positive_prob']:.4f})
    print(fNegative probability: {result['negative_prob']:.4f})
    print("-" * 50)

Note: The code above demonstrates the complete pipeline, but actually fine-tuning BERT requires a GPU and a long time. In production environments, you can also consider using lighter models (such as DistilBERT), or directly use the API of a large model.


Deep Dive into the GPT Series

GPT (Generative Pre-trained Transformer) took a different path from BERT.

Autoregressive Language Model

GPT is "autoregressive": it generates word by word, and at each step uses all previously generated words to predict the next word.

For example, to generate "I like eating apples", the process is:

1. Input "I", predict the next word is "like"

2. Input "I like", predict the next word is "eating"

3. Input "I like eating", predict the next word is "apples"

This approach is naturally suited for text generation.

Decoder-Only Architecture

GPT only uses the Transformer decoder, while BERT only uses the encoder.

The decoder's characteristic is "Masked Self-Attention" - each position can only see positions to its left, not to its right.

This makes sense because text generation is left-to-right; you cannot peek at words that have not been generated yet.

The Discovery of Scaling Laws

The most important discovery of the GPT series is the Scaling Law:Model performance has a power-law relationship with model size, data volume, and compute amount.。

Simply put: the larger the model, the more data, and the longer the training, the better the results. And this improvement is predictable without an obvious plateau.

This is why the GPT series keeps getting "bigger": GPT-1 (117M) → GPT-2 (1.5B) → GPT-3 (175B).


T5 and the Seq2Seq Paradigm

There is also a third paradigm: the Encoder-Decoder structure, represented by T5.

Unified Text-to-Text Framework

The core idea of T5 (Text-to-Text Transfer Transformer) is:Unify all NLP tasks into a "text-to-text" format.。

For example:

Text classification: Input "Classification: This movie is great" → Output "Positive"

Translation: Input "Translate English to Chinese: Hello world" → Output "Hello World"

Summary: Input "Summary: This is an article about..." → Output "This article mainly discusses..."

The benefit of this unified framework is that the same model can handle various tasks without needing to change the model structure for each task.

How the Encoder-Decoder Works

The encoder is responsible for understanding the input text, and the decoder is responsible for generating the output text.

Take translation as an example:

1. The encoder reads "Hello world" and generates an "understood" representation

2. Based on this representation, the decoder generates "Hello World" word by word

The decoder can "attend" to different positions of the encoder—for example, when generating Hello, it pays more attention to "Hello", and when generating "world", it pays more attention to "world".


Comparison of Three Major Paradigms: BERT vs GPT vs T5

Let's compare these three architectures with a table:

FeatureBERTGPTT5
ArchitectureEncoder-onlyDecoder-onlyEncoder-Decoder
AttentionBidirectional (sees both left and right context)Unidirectional (only sees left context)Encoder bidirectional, decoder unidirectional
Pre-training taskMLM (masked language modeling)Autoregressive language modelingSpan Corruption
Excels atUnderstanding tasks: classification, NER, QAGeneration tasks: continuation, dialogueTransformation tasks: translation, summarization, rewriting
Representative modelsBERT、RoBERTa、ALBERTGPT-1/2/3/4、ChatGPTT5、BART、mT5
Typical applicationsSentiment analysis, spam filteringWriting assistant, chatbotMachine translation, document summarization

Today's large language models mostly use a Decoder-only architecture (the GPT route) because it performs best at generation and can be continuously improved through Scaling Laws. However, the other two architectures still have advantages for specific tasks.


Text Classification

Text classification is one of the most common NLP tasks and the best starting task for beginners.

Multiclass Classification in Practice

Multi-class classification means assigning a label to text, and the label has multiple options.

For example: news classification (technology, sports, entertainment, finance), customer service ticket classification, product review classification, etc.

Example

# ============================================
# Text classification in practice: comparison of multiple approaches
# ============================================

from typing import List, Dict, Any

# ============================================
# Approach 1: traditional machine learning + TF-IDF
# Suitable for small datasets, strong interpretability
# ============================================

def traditional_text_classification_demo():
    """Use traditional methods for text classification"""
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.naive_bayes import MultinomialNB
    from sklearn.pipeline import Pipeline

    # Training data
    texts = [
        "I bought a stock today, it went up a lot",
        "Regular fund investment is a good way to manage finances",
        "The forward of this team scored",
        "The NBA playoffs are really exciting",
        "This phone's camera takes great photos",
        "EXAMPLE's programming tutorials are very clear",
        "This movie's plot is very touching",
        "The songs on that album are very nice",
    ]

    labels = ["Finance", "Finance", "Sports", "Sports", "Technology", "Technology", "Entertainment", "Entertainment"]

    # Create pipeline: TF-IDF + Naive Bayes
    pipeline = Pipeline([
        ("tfidf", TfidfVectorizer()),  # Convert text to TF-IDF features
        ("classifier", MultinomialNB()),  # Naive Bayes classifier
    ])

    # Train
    pipeline.fit(texts, labels)

    # Test
    test_texts = [
        "The stock fell, I'm in a bad mood",
        "Basketball game buzzer-beater in the last second",
        "The newly released laptop has strong performance",
        "EXAMPLE's new NLP tutorial",
    ]

    predictions = pipeline.predict(test_texts)
    probabilities = pipeline.predict_proba(test_texts)

    print("Option 1: Traditional machine learning + TF-IDF")
    print("-" * 50)
    for text, pred, probs in zip(test_texts, predictions, probabilities):
        print(f"Text: {text}")
        print(f"Prediction: {pred}")
        print(f"Class probabilities: {dict(zip(pipeline.classes_, probs))}")
        print()

# ============================================
# Option 2: Use Hugging Face pipeline
# Suitable for quick prototyping, no training required
# ============================================

def huggingface_pipeline_demo():
    """Use Hugging Face's pretrained pipeline"""
    print("Option 2: Hugging Face Pipeline")
    print("-" * 50)

    try:
        from transformers import pipeline

        # Use sentiment analysis pipeline (English)
        classifier = pipeline("sentiment-analysis")

        test_texts = [
            "I love EXAMPLE tutorials!",
            "This movie is terrible.",
        ]

        results = classifier(test_texts)

        for text, result in zip(test_texts, results):
            print(f"Text: {text}")
            print(f"Sentiment: {result['label']}")
            print(f"Confidence: {result['score']:.4f}")
            print()

    except Exception as e:
        print(f"Need to install transformers first: pip install transformers")
        print(f"Error: {e}")

# ============================================
# Option 3: Use large model API (recommended)
# Suitable for production environment, best performance
# ============================================

def llm_based_classification_demo():
    """Use large model for text classification (simulation)"""
    print("Option 3: Large model classification (demonstration concept)")
    print("-" * 50)

    # In actual use, call APIs like OpenAI/Anthropic
    # Demonstrate prompt design for classification here

    def classify_with_llm(text: str, categories: List[str]) -> Dict[str, Any]:
        """Simulated function for classification with LLM"""
        # Real scenario: call LLM API
        # prompt = f"""Please classify the following text and only return the category name.
        # Text: {text}
        # Optional categories: {', '.join(categories)}"""

        # Do a simple keyword matching here to simulate
        category_map = {
            "Finance": ["Stock", "Fund", "Wealth Management", "Investment"],
            "Sports": ["Basketball", "Football", "Match", "Goal"],
            "Technology": ["Phone", "Computer", "Programming", "Tutorial"],
            "Entertainment": ["Movie", "Music", "Album", "Song"],
        }

        for category, keywords in category_map.items():
            if any(keyword in text for keyword in keywords):
                return {
                    "category": category,
                    "confidence": 0.9,
                    "reasoning": f"Contains keywords: {', '.join([k for k in keywords if k in text])}"
                }

        return {"category": "Other", "confidence": 0.5}

    test_texts = [
        "Stock went up, I'm very happy",
        "The Football World Cup has opened",
        "EXAMPLE's new programming tutorial",
    ]

    categories = ["Finance", "Sports", "Technology", "Entertainment", "Other"]

    for text in test_texts:
        result = classify_with_llm(text, categories)
        print(f"Text: {text}")
        print(f"Category: {result['category']}")
        print(f"Confidence: {result['confidence']:.4f}")
        if "reasoning" in result:
            print(f"Reason: {result['reasoning']}")
        print()

# ============================================
# Run the demo
# ============================================

if __name__ == "__main__":
    traditional_text_classification_demo()
    print("=" * 50)
    llm_based_classification_demo()

Few-Shot Classification Methods

Few-Shot Learning is a new capability brought by large models—just provide a few examples and the model can learn to classify.

For example:

Text: This product works well → Label: Positive

Text: The quality is too bad → Label: Negative

Text: [your new text] → Label: ?

After seeing these two examples, the large model can understand the task and make reasonable classifications for new text.

Large Model vs Small Model Selection

When to use which approach? Here is a decision guide:

ScenarioRecommended approachRationale
Small data volume (< 1000 entries)Large model few-shot learningSmall models can't learn it; large models generalize well
Large data volume (> 10000 entries)Fine-tune a small model (e.g., BERT)Lower cost, faster inference
Need quick validationLarge model APINo training needed, immediately usable
Production environment, high throughputFine-tune small model + distillationFast inference, controllable cost
Categories change frequentlyLarge model prompt engineeringNo retraining needed, just change the prompt

Named Entity Recognition (NER)

Named Entity Recognition (NER) is the task of identifying specific entities from text, such as person names, place names, and organization names.

Sequence Labeling Framework

NER is usually modeled as a "sequence labeling" task: assign a label to each word indicating whether it is part of an entity.

For example, the sentence: "Zhang SanIn Google 工do", the labels might be:

张三 → B-PER (person name start)

in → O (non-entity)

Google → B-ORG (organization start)

Work → O

BIO Annotation Format

The most common annotation format is BIO:

B-XXX: entity start

I-XXX: inside entity

O: non-entity

For example: "Wang Xiaoming works at Microsoft in Beijing", labeled as:

Wang → B-PER

Xiao → I-PER

Ming → I-PER

at → O

Bei → B-LOC

Jing → I-LOC

's → O

Wei → B-ORG

Ruan → I-ORG

Gong → I-ORG

Si → I-ORG

on → O

class → O

Practical Applications

Let's use Hugging Face for NER:

Example

# ============================================
# Hands-on NER: Using Hugging Face for named entity recognition
# ============================================

def ner_with_huggingface():
    """Use a pretrained NER model"""
    print(NER in action: Hugging Face)
    print("-" * 50)

    try:
        from transformers import pipeline

        # Load NER pipeline
        # English NER
        ner_pipeline = pipeline(
            "ner",
            model="dbmdz/bert-large-cased-finetuned-conll03-english",
            grouped_entities=True  # Merge multiple tokens of the same entity
        )

        test_text = "John Smith works at Google in New York City. He loves EXAMPLE tutorials."

        results = ner_pipeline(test_text)

        print(f"Text: {test_text}")
        print("Recognized entities:")
        for entity in results:
            print(f" - {entity['word']}: {entity['entity_group']} (confidence: {entity['score']:.4f})")
        print()

    except Exception as e:
        print(f"Need to install transformers first: pip install transformers")
        print(f"Error: {e}")

# ============================================
# Chinese NER: using regex + keywords (simple scenario)
# ============================================

def simple_chinese_ner():
    """Simple Chinese NER: using regex and dictionary"""
    import re

    print("Simple Chinese NER: regex + dictionary")
    print("-" * 50)

    # Entity dictionary (would be larger in real scenarios)
    person_names = ["Zhang San", "Li Si", "Wang Xiaoming", "Liu Dehua"]
    organizations = ["Google", "Microsoft", "Alibaba", "Tencent", "EXAMPLE"]
    locations = ["Beijing", "Shanghai", "Shenzhen", "New York", "London"]

    # Regex patterns
    patterns = {
        "PERSON": "|".join(re.escape(name) for name in person_names),
        "ORG": "|".join(re.escape(org) for org in organizations),
        "LOC": "|".join(re.escape(loc) for loc in locations),
        "PHONE": r"1[3-9]\d{9}",  # Phone number
        "EMAIL": r"\w+@\w+\.\w+",  # Email
    }

    def extract_entities(text: str):
        """Extract entities from text"""
        entities = []

        for entity_type, pattern in patterns.items():
            for match in re.finditer(pattern, text):
                entities.append({
                    "text": match.group(),
                    "type": entity_type,
                    "start": match.start(),
                    "end": match.end(),
                })

        # Sort by position
        entities.sort(key=lambda x: x["start"])
        return entities

    # Test
    test_texts = [
        "Zhang San works at Google, phone number is 13800138000",
        "Wang Xiaoming goes to Microsoft in Beijing on a business trip",
        "If you have questions, contact support@example.com",
        "EXAMPLE is a very good learning website",
    ]

    for text in test_texts:
        entities = extract_entities(text)
        print(f"Text: {text}")
        if entities:
            print("Recognized entities:")
            for ent in entities:
                print(f"  - {ent['text']}: {ent['type']}")
        else:
            print("No entities recognized")
        print()

# ============================================
# Large model for NER
# ============================================

def ner_with_llm():
    """Demonstrating the idea of using a large model for NER"""
    print("Large model for NER (demonstration idea)")
    print("-" * 50)

    def extract_entities_with_llm(text: str):
        """Using LLM for NER (simulation)"""
        # Real scenario: construct prompt to call LLM
        # prompt = f"""Please extract named entities from the following text.
        # Only return JSON format, including entity types: PERSON (person name), ORG (organization), LOC (location).
        # Text: {text}"""

        # Here we do a simple simulation
        entities = []

        # Simple keyword matching for simulation
        if "Zhang San" in text:
            entities.append({"text": "Zhang San", "type": "PERSON"})
        if "Google" in text:
            entities.append({"text": "Google", "type": "ORG"})
        if "Beijing" in text:
            entities.append({"text": "Beijing", "type": "LOC"})
        if "EXAMPLE" in text:
            entities.append({"text": "EXAMPLE", "type": "ORG"})

        return entities

    test_text = "Zhang San works at Google in Beijing, and he often uses EXAMPLE to learn."
    entities = extract_entities_with_llm(test_text)

    print(f"Text: {test_text}")
    print("Recognized entities:")
    for ent in entities:
        print(f"  - {ent['text']}: {ent['type']}")

# ============================================
# Run demo
# ============================================

if __name__ == "__main__":
    simple_chinese_ner()
    print("=" * 50)
    ner_with_llm()

Machine Translation

Machine translation is one of the most classic applications of NLP and one of the earliest commercialized applications.

Principles of Neural Machine Translation

Early machine translation was rule-based, then statistical, and now it's all neural.

Neural Machine Translation (NMT) typically uses an Encoder-Decoder architecture:

1. The Encoder reads the source language sentence and generates a semantic representation

2. The Decoder generates the target language sentence based on the semantic representation

The Attention mechanism allows the decoder to "focus" on different positions of the source language sentence when generating each word—this is key to improving translation quality.

Evaluation Metric: BLEU Score

How to measure translation quality? BLEU (Bilingual Evaluation Understudy) is the most commonly used metric.

The core idea of BLEU is:The higher the n-gram overlap between the machine translation result and the human translation result, the better the quality.。

The BLEU score ranges from 0 to 1, the higher the better. Generally:

Below 0.1: basically unintelligible

0.1-0.3: can understand the general meaning

0.3-0.5: good translation quality

Above 0.5: close to human translation

Example

# ============================================
# BLEU score calculation demo
# ============================================

def bleu_score_demo():
    """Demonstrate the calculation of BLEU score"""
    print("BLEU Score Demo")
    print("-" * 50)

    # Note: In practice, it's recommended to use the sacrebleu library
    # Here we do a conceptual demonstration

    from collections import Counter
    import math

    def compute_ngrams(tokens, n):
        """Calculate n-gram"""
        return [tuple(tokens[i:i+n]) for i in range(len(tokens)-n+1)]

    def simple_bleu(reference: str, candidate: str, max_n: int = 4):
        """Simple BLEU calculation (conceptual demonstration)"""
        ref_tokens = reference.split()
        cand_tokens = candidate.split()

        # Calculate brevity penalty
        ref_len = len(ref_tokens)
        cand_len = len(cand_tokens)

        if cand_len == 0:
            return 0.0

        if cand_len <= ref_len:
            bp = math.exp(1 - ref_len / cand_len)
        else:
            bp = 1.0

        # Calculate precision for each n-gram
        precisions = []
        for n in range(1, max_n + 1):
            ref_ngrams = Counter(compute_ngrams(ref_tokens, n))
            cand_ngrams = Counter(compute_ngrams(cand_tokens, n))

            if not cand_ngrams:
                precisions.append(0.0)
                continue

            # Count the number of matches
            matches = 0
            for ngram, count in cand_ngrams.items():
                matches += min(count, ref_ngrams.get(ngram, 0))

            total = sum(cand_ngrams.values())
            precisions.append(matches / total if total > 0 else 0.0)

        # Geometric mean
        if all(p == 0 for p in precisions):
            geo_mean = 0.0
        else:
            # Avoid log(0)
            log_sum = sum(math.log(p + 1e-10) for p in precisions) / max_n
            geo_mean = math.exp(log_sum)

        bleu = bp * geo_mean
        return bleu

    # Test
    reference = "the cat sat on the mat"
    candidates = [
        "the cat sat on the mat",  # Exact match
        "the cat was on the mat",  # One word different
        "the cat sat on a mat",    # Different article
        "a cat sat on the mat",
        "the dog sat on the mat",  # One word error
        "mat the on sat cat the",  # Word order completely scrambled
        "hello world",             # Completely irrelevant
    ]

    print(fReference translation: {reference})
    print()
    for candidate in candidates:
        bleu = simple_bleu(reference, candidate)
        print(fCandidate translation: {candidate})
        print(fBLEU score: {bleu:.4f})
        print()

    # Chinese example
    print(Chinese translation example:)
    print("-" * 30)
    ref = I like to eat apples
    cand1 = I like to eat apples
    cand2 = I love to eat apples
    cand3 = Apples I like to eat

    print(fReference: {ref})
    print(fCandidate 1: {cand1}, BLEU: {simple_bleu(ref, cand1):.4f})
    print(fCandidate 2: {cand2}, BLEU: {simple_bleu(ref, cand2):.4f})
    print(fCandidate 3: {cand3}, BLEU: {simple_bleu(ref, cand3):.4f})


if __name__ == "__main__":
    bleu_score_demo()

Translation Quality in the LLM Era

After the advent of large models, machine translation quality has reached a new level.

Traditional translation models are trained on bilingual parallel data, while large models are trained on massive amounts of text and have a deeper understanding of language.

Advantages of large models in translation:

1. Better context understanding: can handle ambiguity and polysemy

2. Style control: can specify the tone and style of the translation

3. Terminology consistency: can provide a glossary to ensure consistent translation of proper nouns

4. Multilingual capability: one model can translate between dozens or hundreds of languages


Sentiment Analysis

Sentiment analysis is determining the emotional tendency of a text—positive, negative, or neutral.

Fine-Grained Sentiment Analysis

Simple sentiment analysis is binary classification (positive/negative); more fine-grained can be:

1. Sentiment scoring: 1-5 stars, 1 is the most negative, 5 is the most positive

2. Sentiment classification: anger, joy, sadness, surprise, etc.

3. Aspect-Based Sentiment Analysis: not only judges overall sentiment but also analyzes sentiment toward specific aspects.

For example: "The food at this restaurant is delicious, but the service is too slow."

Overall sentiment: neutral to positive

Aspect-level sentiment:

Food → Positive

Service → Negative

Multilingual Sentiment Models

Current sentiment models can handle multiple languages without needing separate training for each language.

Common multilingual models:

Large models such as XLM-RoBERTa, mBERT (multilingual BERT), Llama, and Claude.


Text Summarization

Text summarization is the process of shortening long text while preserving core information.

Extractive vs Abstractive Summarization

There are two main summarization approaches:

Extractive summarization: Select the most important sentences from the original text and put them together.

Advantages: information is accurate and will not fabricate

Disadvantages: not fluent enough, may contain redundant information

Abstractive summarization: Let the model "understand" the original text, then write a summary in its own words.

Advantages: fluent, concise, and can summarize the original

Disadvantages: may "hallucinate", fabricating information not present in the original text

In the era of large models, abstractive summarization has become mainstream, but attention must be paid to the hallucination problem when using it.

ROUGE Evaluation Metric

How is summarization quality evaluated? ROUGE is the most commonly used metric.

ROUGE has several variants:

ROUGE-1: matching based on unigrams (single words)

ROUGE-2: matching based on bigrams (two words)

ROUGE-L: based on the longest common subsequence

Similar to BLEU, ROUGE also measures the overlap between machine-generated summaries and human reference summaries.

Example

# ============================================
# Text summarization: ROUGE metrics + simple extractive summarization
# ============================================

def rouge_demo():
    """ROUGE metrics demo"""
    print("ROUGE metrics demo")
    print("-" * 50)

    from collections import Counter

    def compute_rouge_1(reference: str, candidate: str) -> float:
        """Compute ROUGE-1 (unigram F1)"""
        ref_tokens = reference.split()
        cand_tokens = candidate.split()

        if not ref_tokens or not cand_tokens:
            return 0.0

        ref_counts = Counter(ref_tokens)
        cand_counts = Counter(cand_tokens)

        # Calculate the number of matches
        matches = 0
        for token, count in cand_counts.items():
            matches += min(count, ref_counts.get(token, 0))

        precision = matches / len(cand_tokens)
        recall = matches / len(ref_tokens)

        if precision + recall == 0:
            return 0.0

        f1 = 2 * precision * recall / (precision + recall)
        return f1

    def compute_lcs(a: list, b: list) -> int:
        """Compute the length of the longest common subsequence"""
        m, n = len(a), len(b)
        dp = [[0] * (n + 1) for _ in range(m + 1)]

        for i in range(1, m + 1):
            for j in range(1, n + 1):
                if a[i-1] == b[j-1]:
                    dp[i][j] = dp[i-1][j-1] + 1
                else:
                    dp[i][j] = max(dp[i-1][j], dp[i][j-1])

        return dp[m][n]

    def compute_rouge_l(reference: str, candidate: str) -> float:
        """Compute ROUGE-L (longest common subsequence F1)"""
        ref_tokens = reference.split()
        cand_tokens = candidate.split()

        if not ref_tokens or not cand_tokens:
            return 0.0

        lcs_len = compute_lcs(ref_tokens, cand_tokens)

        precision = lcs_len / len(cand_tokens)
        recall = lcs_len / len(ref_tokens)

        if precision + recall == 0:
            return 0.0

        f1 = 2 * precision * recall / (precision + recall)
        return f1

    # Test
    reference = "Natural language processing is an important branch of artificial intelligence, studying the interaction between computers and human language."
    candidates = [
        "Natural language processing is an important branch of artificial intelligence, studying the interaction between computers and human language.",  # Exact match
        "Natural language processing studies the interaction between computers and human language, and is an important branch of artificial intelligence.",  # Order changed
        "Natural language processing is a branch of artificial intelligence, studying the interaction between computers and language.",  # Abridged version
        "EXAMPLE is a programming learning website.",  # Irrelevant
    ]

    print(f"Reference summary: {reference}")
    print()
    for candidate in candidates:
        rouge1 = compute_rouge_1(reference, candidate)
        rougel = compute_rouge_l(reference, candidate)
        print(f"Candidate summary: {candidate}")
        print(f"ROUGE-1: {rouge1:.4f}")
        print(f"ROUGE-L: {rougel:.4f}")
        print()


def simple_extractive_summarization():
    Simple extractive summarization: selecting important sentences based on word frequency
    print(Simple extractive summarization demo)
    print("-" * 50)

    import re
    from collections import Counter

    def split_sentences(text: str) -> list:
        Simple Chinese sentence segmentation
        # Split sentences by periods, exclamation marks, and question marks
        sentences = re.split(r'[。!?]', text)
        # Filter out empty strings
        sentences = [s.strip() for s in sentences if s.strip()]
        return sentences

    def get_word_frequency(text: str) -> Counter:
        """Calculate word frequency (simple demo: simulated with character-level n-grams)"""
        # In real scenarios, a word segmentation tool such as jieba should be used
        # Here we simply use uni-gram
        chars = [c for c in text if c.strip()]
        return Counter(chars)

    def sentence_score(sentence: str, word_freq: Counter) -> float:
        """Calculate sentence scores"""
        if not sentence:
            return 0.0

        # Simple score: sum of word frequencies in the sentence / sentence length
        total_score = 0.0
        for char in sentence:
            total_score += word_freq.get(char, 0)

        # Normalization
        return total_score / (len(sentence) + 1)

    def extractive_summary(text: str, num_sentences: int = 2) -> str:
        """Extractive summarization"""
        sentences = split_sentences(text)

        if len(sentences) <= num_sentences:
            return text

        # Calculate word frequency
        word_freq = get_word_frequency(text)

        # Calculate the score for each sentence
        scores = [(i, sentence_score(s, word_freq)) for i, s in enumerate(sentences)]

        # Sort by score and select the top ones
        scores.sort(key=lambda x: x[1], reverse=True)
        top_indices = [i for i, score in scores[:num_sentences]]
        top_indices.sort()  # Keep the original order

        # Concatenate the results
        summary = "。".join(sentences[i] for i in top_indices) + "。"
        return summary

    # Test text
    test_text = """
Natural language processing is an important branch of artificial intelligence. It studies how to make computers understand and generate human language.
This technology has many applications, including machine translation, sentiment analysis, question-answering systems, and more.
In recent years, the emergence of large language models has brought major breakthroughs in natural language processing.
The EXAMPLE website provides many excellent programming tutorials to help developers learn new skills.
"""


    print("Original text:")
    print(test_text.strip())
    print()

    summary = extractive_summary(test_text, num_sentences=2)
    print("Extractive summary:")
    print(summary)


if __name__ == "__main__":
    rouge_demo()
    print("=" * 50)
    simple_extractive_summarization()
Other extensions