Deep Learning Basics

Deep learning is not some mysterious black technology; it is essentially a mathematical function fitter—give it large amounts of input-output data, and it automatically finds the complex relationships between them.

For example: you give it a million pictures of cats, with each picture corresponding to the label "cat". Deep learning automatically figures out what combinations of pixels look like a cat.

Another example: you give it a billion sentences of human language, with the next word of each sentence as the target. Deep learning automatically learns, given the preceding context, what the next word is most likely to be.

The core advantage of deep learning is that as data volume increases, its performance continues to improve.This is something traditional machine learning cannot do.

In this chapter, we will start from scratch and break down every component of deep learning step by step:

  • Mathematical Foundations of Neural Networks
  • The Computational Process of Forward Propagation
  • The Role of Activation Functions
  • Loss Function Design
  • The Principle of Backpropagation
  • Optimizer Selection
  • Regularization Techniques
  • Normalization Methods
  • Learning Rate Scheduling

Finally, we implement a complete multi-layer perceptron (MLP) with PyTorch, connecting all the knowledge points together.

This chapter involves some mathematics, but don't worry—we will explain every formula intuitively rather than asking you to memorize it. The focus is on understanding why it is designed this way, not how to calculate.


Mathematical Foundations of Neural Networks

Deep learning is built on three branches of mathematics: linear algebra, calculus, and probability theory.

You don't need to be a math expert, but you do need to understand a few core concepts.

Linear Algebra: Matrix Multiplication

The most basic operation in neural networks is matrix multiplication.

Why use matrices? Because they can concisely represent the process of "multiple inputs being transformed by multiple neurons."

Let's first look at an intuitive example:

Example

# ============================================
# Intuitive Understanding of Matrix Multiplication
# ============================================

import numpy as np

# Assume the input is a 3-dimensional vector: [height, weight, age]
# Units are: centimeters, kilograms, years
x = np.array([175, 70, 25])  # Input vector, shape (3,)
print(f"Input x: {x}")
print(f"Input shape: {x.shape}")

# Weight matrix W: 2 neurons, each neuron receives 3 inputs
# Shape is (output dimension, input dimension) = (2, 3)
W = np.array([
    [0.1, 0.2, 0.3],   # Weights of the 1st neuron
    [0.4, 0.5, 0.6],   # Weights of the 2nd neuron
])
print(f"\nWeight matrix W:\n{W}")
print(f"Weight shape: {W.shape}")

# Bias b: one bias per neuron
b = np.array([0.1, 0.2])  # Shape (2,)
print(f"\nBias b: {b}")
print(f"Bias shape: {b.shape}")

# Matrix multiplication: y = W · x + b
# Note: numpy's @ operator represents matrix multiplication
y = W @ x + b
print(f"\nOutput y: {y}")
print(f"Output shape: {y.shape}")

# Let's manually verify
y1 = 0.1 * 175 + 0.2 * 70 + 0.3 * 25 + 0.1  # The 1st neuron
y2 = 0.4 * 175 + 0.5 * 70 + 0.6 * 25 + 0.2  # The 2nd neuron
print(f"\nManual calculation verification: [{y1}, {y2}]")

The geometric meaning of matrix multiplication is a "linear transformation"—it can rotate, scale, and stretch vector spaces.

But linear transformations alone are not enough.

If every layer is just matrix multiplication, then no matter how deep the network is, the entire network is equivalent to a single-layer network, because the composition of multiple linear transformations is still a linear transformation.

This is why we need activation functions—to introduce nonlinearity.

Calculus: Derivatives and the Chain Rule

The core of training neural networks is "gradient descent"—finding the direction of parameters that minimizes the loss.

The gradient is the "derivative of a multivariable function"—it tells us how much the loss changes when each parameter changes a little.

Let's look at a simple example:

Example

# ============================================
# Intuitive understanding of derivatives
# ============================================

def f(x):
    """A simple function: f(x) = x²"""
    return x ** 2

def numerical_derivative(f, x, h=1e-6):
    """Numerical derivative: use a small increment h to approximate the derivative
Definition of derivative: f'(x) = lim(h→0) [f(x+h) - f(x)] / h
    """

    return (f(x + h) - f(x)) / h

# Compute the derivative at x=3
x = 3.0
df_dx = numerical_derivative(f, x)
print(f"f({x}) = {f(x)}")
print(f"f'({x}) ≈ {df_dx}")
print(f"Analytical solution (exact value): 2 * {x} = {2 * x}")

# Understand the meaning of the derivative:
# The derivative 6 means: at x=3, if x increases by 1, f(x) increases by about 6
x_new = x + 0.01
f_new = f(x_new)
print(f"\nx increases from {x} to {x_new}")
print(f"f(x) changes from {f(x)} to {f_new}")
print(f"Actual increase: {f_new - f(x)}")
print(f"Derivative prediction: {df_dx * 0.01}")

For multivariable functions, we need to compute the partial derivative with respect to each variable, then combine them into a vector—this is the gradient.

The chain rule is the mathematical foundation of backpropagation. It allows us to "break down layer by layer" the derivatives of composite functions.

Example

# ============================================
# Intuitive understanding of the chain rule
# ============================================

# Consider a composite function: y = f(g(x))
# Where: g(x) = x², f(z) = z³
# Then: y = (x²)³ = x⁶

def g(x):
    return x ** 2

def f(z):
    return z ** 3

def y(x):
    return f(g(x))

# Manually compute the derivative (chain rule):
# dy/dx = df/dz * dz/dx = 3*z² * 2*x = 3*(x²)² * 2*x = 6*x⁵

x = 2.0
z = g(x)  # z = 4
dy_dx_chain = 3 * (z ** 2) * 2 * x  # Chain rule calculation
print(f"Chain rule calculation: dy/dx = {dy_dx_chain}")

# Numerical derivative verification
def numerical_derivative_y(x, h=1e-6):
    return (y(x + h) - y(x)) / h

dy_dx_numerical = numerical_derivative_y(x)
print(f"Numerical derivative verification: dy/dx ≈ {dy_dx_numerical}")

# Analytical solution: 6*x^5 = 6*32 = 192
print(f"Analytical solution: 6 * {x}^5 = {6 * (x ** 5)}")

The core idea of the chain rule is:Break a complex function into simple functions, differentiate each, then multiply them together.。

Backpropagation is the application of the chain rule in neural networks—starting from the output, it propagates gradients backward layer by layer.

Probability Theory: Conditional Probability

In classification problems, neural networks often output a "probability distribution"—given an input, what is the probability of each class.

Conditional probability P(Y|X) represents "the probability of Y occurring given that X has occurred."

:.2%}

:.2%}

Example

# ============================================
# Use a neural network to output a probability distribution
# ============================================

import numpy as np

def softmax(x):
    """Softmax function: converts arbitrary real numbers into a probability distribution
The output values sum to 1, and each value is between [0, 1]
    """

    # Subtract the maximum value to prevent numerical overflow
    exp_x = np.exp(x - np.max(x))
    return exp_x / np.sum(exp_x)

# Suppose the last layer of the neural network outputs 3 "scores" (logits)
# They correspond to "cat", "dog", "bird" respectively
logits = np.array([2.0, 1.0, 0.5])
print(fNeural network output (logits): {logits})

# Use Softmax to convert to probabilities
probabilities = softmax(logits)
print(fProbability distribution: {probabilities})
print(fSum of probabilities: {np.sum(probabilities)})

# Interpretation:
# P(cat|input image) ≈ 0.67
# P(dog|input image) ≈ 0.24
# P(bird|input image) ≈ 0.09
print("\nCategory probabilities:)
print(fCat: {probabilities)
print(fDog: {probabilities)
print(fBird: {probabilities)

Core points of three mathematical foundations:

Mathematical branchCore conceptUse in deep learning
Linear algebraMatrix multiplication, vectors, tensorsRepresent network structure, efficiently compute forward propagation
CalculusDerivatives, chain rule, gradientsBackpropagation, update parameters
Probability theoryConditional probability, probability distributionModel uncertainty, design loss functions

Forward Propagation

Forward propagation is the computation process where input data "flows through" the neural network, from the input layer to the hidden layer to the output layer.

Computation Graph

Representing a neural network as a computation graph makes the logic of forward propagation and backpropagation clearer.

A computation graph consists of nodes (operations) and edges (data flow).

Computation graph of a simple two-layer neural network:

Example

# ============================================
# Use a computation graph to understand forward propagation
# ============================================

import numpy as np

def relu(x):
    """ReLU activation function: max(0, x)"""
    return np.maximum(0, x)

# A simple two-layer neural network
# Input layer (2) → hidden layer (3) → output layer (2)

# Input: x = [x1, x2]
x = np.array([1.0, 2.0])
print(fInput x: {x})

# First layer: input → hidden layer
W1 = np.array([
    [0.1, 0.2],  # 1st hidden neuron
    [0.3, 0.4],  # 2nd hidden neuron
    [0.5, 0.6],  # 3rd hidden neuron
])
b1 = np.array([0.1, 0.2, 0.3])

# Second layer: hidden layer → output layer
W2 = np.array([
    [0.1, 0.2, 0.3],  # 1st output neuron
    [0.4, 0.5, 0.6],  # 2nd output neuron
])
b2 = np.array([0.1, 0.2])

# Forward propagation computation graph
# Step 1: hidden layer linear transformation
z1 = W1 @ x + b1
print(f"\nHidden layer linear transformation z1: {z1})

# Step 2: hidden layer activation function
a1 = relu(z1)
print(fHidden layer activation a1: {a1})

# Step 3: output layer linear transformation
z2 = W2 @ a1 + b2
print(fOutput layer linear transformation z2: {z2})

# Step 4: output layer activation (for classification problems, usually Softmax)
a2 = softmax(z2)  # Use the softmax function defined earlier
print(fOutput layer activation a2: {a2})

The advantage of the computation graph is that every step is clear, and during backpropagation you just compute gradients in the reverse order.

Neural Networks from the Perspective of Matrix Multiplication

From the perspective of matrix multiplication, a neural network is a series of "linear transformations + nonlinear activations".

Each layer can be abstracted as:

aᵢ = activation(Wᵢ · aᵢ₋₁ + bᵢ)

where a₀ is the input x.

This formula is concise yet powerful—it can represent almost all neural network structures, from simple logistic regression to complex Transformers.

Note the dimension matching for matrix multiplication: if the input vector is n-dimensional and the next layer has m neurons, then the weight matrix W must have shape m×n. This way, after W·x, you get an m-dimensional vector.

The Role of Activation Functions

As mentioned earlier: without activation functions, a deep network is equivalent to a single-layer network.

Use a simple example to prove this:

Example

# ============================================
# Why activation functions are needed
# ============================================

import numpy as np

# Suppose there is a two-layer network, but neither layer has an activation function

# Input
x = np.array([1.0, 2.0])

# First layer weights and bias
W1 = np.array([[0.1, 0.2], [0.3, 0.4]])
b1 = np.array([0.1, 0.2])

# Second layer weights and bias
W2 = np.array([[0.5, 0.6], [0.7, 0.8]])
b2 = np.array([0.3, 0.4])

# Compute the two layers separately
z1 = W1 @ x + b1
z2 = W2 @ z1 + b2
print(f"Two-layer calculation results: {z2}")

# Merge into a single layer for calculation
# It can be proved mathematically: z2 = (W2·W1)·x + (W2·b1 + b2)
W_combined = W2 @ W1
b_combined = W2 @ b1 + b2
z_combined = W_combined @ x + b_combined
print(f"Merged into one-layer calculation result: {z_combined}")

# The results are exactly the same!
print(f"\n"The two results are equal: {np.allclose(z2, z_combined)}")
print("Conclusion: without activation functions, deep network = single-layer network")

Activation functions are the key component that makes deep networks meaningful—they introduce nonlinearity and allow the network to learn complex patterns.


Activation Functions in Detail

An activation function determines how a neuron's output responds to its input.

A good activation function should satisfy several conditions:

  • Nonlinearity: this is a must
  • Differentiability: so gradients can be computed
  • Simple computation: both forward and backward propagation should be fast
  • No saturation: gradients will not vanish or explode

Let's look at the most commonly used activation functions.

Sigmoid: Saturation and Vanishing Gradients

Sigmoid is one of the earliest activation functions. It compresses any real number to between (0, 1).

Formula: σ(x) = 1 / (1 + e⁻ˣ)

Example

# ============================================
# Sigmoid activation function
# ============================================

import numpy as np
import matplotlib.pyplot as plt

def sigmoid(x):
    """Sigmoid function"""
    return 1 / (1 + np.exp(-x))

def sigmoid_derivative(x):
    """Derivative of Sigmoid: σ'(x) = σ(x) * (1 - σ(x))"""
    s = sigmoid(x)
    return s * (1 - s)

# Test some values
x_values = [-10, -5, -2, 0, 2, 5, 10]
print("x    | sigmoid(x) | sigmoid'(x)")
print("-" * 40)
for x in x_values:
    s = sigmoid(x)
    d = sigmoid_derivative(x)
    print(f"{x:4} | {s:10.6f} | {d:12.6f}")

# Observation: when x is very large or very small, the derivative approaches 0
# This is the "vanishing gradient" problem
x_big = 10.0
print(f"\n"When x = {x_big}:")
print(f" sigmoid(x) = {sigmoid(x_big):.10f} (nearly 1)")
print(f" sigmoid'(x) = {sigmoid_derivative(x_big):.10f} (nearly 0)")
print("The gradient vanishes! During backpropagation, the gradient cannot be passed back.")

The problems with Sigmoid are obvious:

  • When |x| > 6, the gradient is almost 0, causing vanishing gradients
  • The output is not centered at 0, which affects optimization
  • Computing the exponential is relatively slow

Due to these problems, Sigmoid is rarely used in hidden layers in modern deep learning.

ReLU and Variants

ReLU (Rectified Linear Unit) is the most commonly used activation function today.

Formula: ReLU(x) = max(0, x)

Simple but very effective.

Example

# ============================================
# ReLU and its variants
# ============================================

import numpy as np

def relu(x):
    """ReLU:max(0, x)"""
    return np.maximum(0, x)

def relu_derivative(x):
    """Derivative of ReLU: 1 when x > 0, otherwise 0"""
    return (x > 0).astype(float)

def leaky_relu(x, alpha=0.01):
    """Leaky ReLU: x when x > 0, otherwise alpha*x"""
    return np.where(x > 0, x, alpha * x)

def gelu(x):
    """GELU:Gaussian Error Linear Unit
This is the activation function commonly used in Transformer
Approximation formula: 0.5 * x * (1 + tanh(sqrt(2/pi) * (x + 0.044715*x^3)))
    """

    cdf = 0.5 * (1.0 + np.tanh(np.sqrt(2 / np.pi) * (x + 0.044715 * np.power(x, 3))))
    return x * cdf

# Test
x_values = [-5, -2, -1, 0, 1, 2, 5]
print("x    | ReLU | Leaky | GELU")
print("-" * 45)
for x in x_values:
    r = relu(x)
    lr = leaky_relu(x)
    g = gelu(x)
    print(f"{x:4} | {r:4.2f} | {lr:5.2f} | {g:5.2f}")

# Advantages of ReLU:
print("\n"Advantages of ReLU:")
print("1. Simple computation: only needs the max operation")
print("2. No saturation (positive interval): the gradient is always 1")
print("3. Sparsity: some neuron outputs are 0, increasing model robustness")

# Problem with ReLU: "dying ReLU"
print("\n"Problem with ReLU: dying ReLU")
print(If a neuron's input is always negative, then:)
print(- Output is always 0)
print(- Gradient is always 0)
print(- Parameters are never updated)
print(- This neuron is 'dead')

Comparison of several common activation functions:

Activation functionFormulaAdvantagesDisadvantagesApplicable scenarios
Sigmoid1/(1+e⁻ˣ)Output in (0,1), interpretable as probabilityVanishing gradient, non-zero-centered, slow computationBinary classification output layer (rarely used for hidden layers)
ReLUmax(0, x)Fast computation, non-saturation, sparsityDead ReLU, non-zero-centered outputMost hidden layers (default choice)
Leaky ReLUmax(αx, x)Solves the dead ReLU problemα needs to be selected manuallyAlternative when ReLU fails
GELUx·Φ(x)Smooth, works well in TransformerSlightly more complex computationTransformer, BERT, GPT, etc.
Swish/SwiGLUx·σ(βx)Superior to ReLU in deep networksComplex computationLarge models such as PaLM, LLaMA

Selection Principles

Recommended choices for activation functions:

  • Hidden layer: Prefer ReLU, simple and effective
  • If you encounter dead ReLU: Try Leaky ReLU or GELU
  • Transformer: GELU or SwiGLU is the standard choice
  • Binary classification output layer: Sigmoid (outputs probability)
  • Multi-class classification output layer: Softmax (outputs probability distribution)
  • Regression output layer: No activation (outputs any real number)

Don't overthink the choice of activation function. Usually ReLU is good enough. When you confirm ReLU is the bottleneck, it's not too late to switch to others.


Loss Functions

The loss function measures the gap between the "model's prediction" and the "true answer".

The goal of training is to make this gap as small as possible.

Cross-Entropy Loss (Classification)

The most commonly used loss function for classification problems is Cross-Entropy Loss.

Its intuitive meaning is: "how different is the probability distribution predicted by the model from the true distribution."

Example

# ============================================
# Cross-Entropy Loss
# ============================================

import numpy as np

def cross_entropy_loss(probs, target_one_hot):
    """Cross-Entropy Loss
probs: model-predicted probability distribution, shape (batch_size, num_classes)
target_one_hot: one-hot encoding of true labels, same shape as above
    """

    # Add epsilon to prevent log(0)
    epsilon = 1e-12
    probs = np.clip(probs, epsilon, 1.0 - epsilon)
    # cross_entropy = -sum(target * log(pred))
    return -np.sum(target_one_hot * np.log(probs)) / len(probs)

def softmax(x):
    """Softmax function"""
    exp_x = np.exp(x - np.max(x, axis=-1, keepdims=True))
    return exp_x / np.sum(exp_x, axis=-1, keepdims=True)

# Example: three-class classification problem, classes are "cat", "dog", "bird"

# Assume the true label is "cat" (index 0), one-hot encoded
target = np.array([[1, 0, 0]])  # The correct answer is class 0

# Scenario 1: The model confidently predicts "cat"
logits1 = np.array([[5.0, 0.5, 0.1]])  # Cat has the highest score
probs1 = softmax(logits1)
loss1 = cross_entropy_loss(probs1, target)
print("Scenario 1: Model confidently predicts correctly")
print(f" Probability distribution: {probs1)
print(f" Loss: {loss1:.6f} (very small)")

# Scenario 2: The model is uncertain
logits2 = np.array([[1.0, 0.9, 0.8]])  # The scores of the three classes are similar
probs2 = softmax(logits2)
loss2 = cross_entropy_loss(probs2, target)
print("\n"Scenario 2: Model is uncertain")
print(f" Probability distribution: {probs2)
print(f" Loss: {loss2:.6f} (medium)")

# Scenario 3: The model predicts incorrectly
logits3 = np.array([[0.1, 5.0, 0.5]])  # Dog has the highest score
probs3 = softmax(logits3)
loss3 = cross_entropy_loss(probs3, target)
print("\n"Scenario 3: Model predicts incorrectly")
print(f" Probability distribution: {probs3)
print(f" Loss: {loss3:.6f} (very large)")

A simplified version of cross-entropy loss: when labels are class indices instead of one-hot encoding, a more efficient computation method can be used:

loss = -log(probs[target_class])

This is why in PyTorch, CrossEntropyLoss directly accepts class indices as labels.

MSE Loss (Regression)

For regression problems (predicting continuous values), the most commonly used loss is Mean Squared Error (MSE).

Formula: MSE = mean((pred - target)²)

Example

# ============================================
# MSE Loss
# ============================================

import numpy as np

def mse_loss(pred, target):
    """Mean Squared Error Loss"""
    return np.mean((pred - target) ** 2)

def mae_loss(pred, target):
    """Mean Absolute Error Loss"""
    return np.mean(np.abs(pred - target))

# Example: predicting house prices (unit: ten thousand yuan)

# Actual house prices
target = np.array([100, 200, 150])

# Scenario 1: Prediction is very accurate
pred1 = np.array([102, 195, 148])
mse1 = mse_loss(pred1, target)
mae1 = mae_loss(pred1, target)
print("Scenario 1: Accurate prediction")
print(f" Prediction: {pred1}")
print(f" Actual: {target}")
print(f"  MSE: {mse1:.2f}")
print(f"  MAE: {mae1:.2f}")

# Scenario 2: Prediction has a large error
pred2 = np.array([102, 280, 148])  # The second house prediction is 800,000 higher
mse2 = mse_loss(pred2, target)
mae2 = mae_loss(pred2, target)
print("\n"Scenario 2: There is a large error")
print(f" Prediction: {pred2}")
print(f" Actual: {target}")
print(f" MSE: {mse2:.2f} (greatly amplified by the large error)")
print(f"  MAE: {mae2:.2f}")

# MSE vs MAE
print("\nMSE vs MAE:")
print("- MSE is more sensitive to large errors (because of squaring)")
print("- MAE is more robust to outliers")
print("- Generally, MSE is used more often")

Next-Word Prediction Loss for Language Models

The training objective of large language models (e.g., GPT) is simple: given the preceding context, predict the next word.

This is essentially a multi-class classification problem—each word in the vocabulary is a class.

Example

# ============================================
# Loss function for language models
# ============================================

import numpy as np

def language_model_loss(logits, targets, vocab_size):
    """Language model loss
Essentially, it is the cross-entropy loss at each position.
    """

    batch_size, seq_len, _ = logits.shape
    total_loss = 0.0

    # Compute the loss at each position in the sequence
    for i in range(seq_len):
        # Predicted scores at this position
        step_logits = logits[:, i, :]
        # Convert to probabilities
        step_probs = softmax(step_logits)
        # Index of the target word
        step_target = targets[:, i]
        # Take only the probability of the target word
        for b in range(batch_size):
            target_idx = step_target[b]
            target_prob = step_probs[b, target_idx]
            # Loss = -log(p)
            total_loss += -np.log(target_prob + 1e-12)

    return total_loss / (batch_size * seq_len)

# A simple example
vocab_size = 10000  # Assume the vocabulary size is 10,000
batch_size = 2
seq_len = 3

# Model output: (batch_size, seq_len, vocab_size)
logits = np.random.randn(batch_size, seq_len, vocab_size)

# True targets: the word index that should be predicted at each position
targets = np.array([
    [123, 456, 789],   # Target word for the first sequence
    [234, 567, 890],   # Target word for the second sequence
])

loss = language_model_loss(logits, targets, vocab_size)
print(f"Language model loss: {loss:.4f}")
print("\n"The meaning of this loss:")
print("- For each position in the sequence")
print("- Predict the next word")
print("- Compute the average cross-entropy over all positions")

Summary of common loss functions:

Task typeLoss functionOutput layer activationLabel format
Binary classificationBinary Cross-EntropySigmoid0 or 1
Multi-class classificationCross-EntropySoftmaxClass index
Multi-label classificationBinary Cross-EntropySigmoidMultiple 0/1
RegressionMSE / MAENoneContinuous value
Language modelCross-EntropySoftmaxWord index

Backpropagation

Backpropagation is the core algorithm for training neural networks.

Its goal is to compute the gradient of the loss function with respect to each parameter, and then use gradient descent to update the parameters.

Derivation via the Chain Rule

We will use a simple example to derive backpropagation step by step.

Example

# ============================================
# Manually implement backpropagation
# ============================================

import numpy as np

def relu(x):
    return np.maximum(0, x)

def relu_derivative(x):
    return (x > 0).astype(float)

def mse_loss(pred, target):
    return np.mean((pred - target) ** 2)

# The simplest neural network: one linear transformation
# y = w * x + b
# Loss = (y - target)^2

# Input and target
x = np.array(2.0)
target = np.array(7.0)

# Initialize parameters
w = np.array(1.0)
b = np.array(0.0)

print(f"Initial parameters: w = {w}, b = {b}")
print(f"Target: {target}")

# ============================================
# Step 1: Forward propagation
# ============================================
y = w * x + b
loss = (y - target) ** 2
print(f"\n"Forward propagation:")
print(f"  y = {y}")
print(f"  loss = {loss}")

# ============================================
# Step 2: Backpropagation (manually compute gradients)
# ============================================
# We need to compute: d_loss/d_w, d_loss/d_b

# Use the chain rule step by step:
# loss = (y - target)^2
# d_loss/d_y = 2 * (y - target)
d_loss_d_y = 2 * (y - target)

# y = w * x + b
# d_y/d_w = x
# d_y/d_b = 1
d_y_d_w = x
d_y_d_b = 1.0

# Chain rule:
# d_loss/d_w = d_loss/d_y * d_y/d_w
# d_loss/d_b = d_loss/d_y * d_y/d_b
d_loss_d_w = d_loss_d_y * d_y_d_w
d_loss_d_b = d_loss_d_y * d_y_d_b

print(f"\n"Backpropagation:")
print(f"  d_loss/d_y = {d_loss_d_y}")
print(f"  d_loss/d_w = {d_loss_d_w}")
print(f"  d_loss/d_b = {d_loss_d_b}")

# ============================================
# Step 3: Update parameters using gradient descent
# ============================================
learning_rate = 0.1
w_new = w - learning_rate * d_loss_d_w
b_new = b - learning_rate * d_loss_d_b

print(f"\n"Parameter update:")
print(f"  w: {w} -> {w_new}")
print(f"  b: {b} -> {b_new}")

# Verify: useNewParameterbeforetoward传播 # Verification: Forward propagation with new parameters
y_new = w_new * x + b_new
loss_new = (y_new - target) ** 2
print(f"\nVerification: Verification:)
print(f"  旧预测: {y}, 旧损失: {loss}" " Old prediction: {y}, old loss: {loss}")
print(f"  New预测: {y_new}, New损失: {loss_new}" " New prediction: {y_new}, new loss: {loss_new}")
print(f" Loss decreased!" " Loss decreased!")

Although this example is simple, it contains all the core ideas of backpropagation:

  1. beforetoward传播:ComputeAllmiddleVariable Forward propagation: compute all intermediate variables
  2. 反toward传播: fromOutputstart,use链formula法ThenlayerlayerCompute梯degree Backpropagation: starting from the output, use the chain rule to compute gradients layer by layer
  3. ParameterUpdate:use梯degreeBottom降UpdateParameter Parameter update: update parameters using gradient descent

Gradient Flow in Computation Graphs

formulti-layer神经Networking,梯degreeWhen computingGraphMedium反towardflow动。 For multi-layer neural networks, gradients flow backward through the computational graph.

everyonelayerof梯degreeDependencyinafteronelayer传comeof梯degree。 The gradient of each layer depends on the gradient passed from the next layer.

Example

# ============================================
# twolayerNetworkingof反toward传播 # Backpropagation for a two-layer network
# ============================================

import numpy as np

def relu(x):
    return np.maximum(0, x)

def relu_derivative(x):
    return (x > 0).astype(float)

# A two-layer network # A two-layer network
# Input (2) → Hidden藏layer (3) → Output (1) # Input (2) → Hidden layer (3) → Output (1)

# Data # Data
x = np.array([1.0, 2.0])  # Input # Input
target = np.array([5.0])  # Target output # Target output

# Parameter initialization # Parameter initialization
W1 = np.array([[0.1, 0.2], [0.3, 0.4], [0.5, 0.6]])  # (3, 2)
b1 = np.array([0.1, 0.2, 0.3])  # (3,)
W2 = np.array([[0.7, 0.8, 0.9]])  # (1, 3)
b2 = np.array([0.4])  # (1,)

print(Initial parameters have been set.)

# ============================================
# Forward propagation # Forward propagation
# ============================================
print("\n=== Forward propagation ===" === Forward propagation ===)

# First layer # First layer
z1 = W1 @ x + b1  # (3,)
a1 = relu(z1)     # (3,)
print(f"z1 = {z1}")
print(f"a1 = {a1}")

# Second layer # Second layer
z2 = W2 @ a1 + b2  # (1,)
a2 = z2  # RegressionTask, outputlayerNoaddactivate # Regression task, no activation on the output layer
print(f"z2 = {z2}")
print(f"a2 = {a2}")

# Loss # Loss
loss = np.mean((a2 - target) ** 2)
print(f"loss = {loss}")

# ============================================
# Backpropagation # Backpropagation
# ============================================
print("\n=== Backpropagation ===" === Backpropagation ===)

# Output layer gradient # Output layer gradient
d_loss_d_a2 = 2 * (a2 - target)  # (1,)
d_a2_d_z2 = 1.0  # No activation function # No activation function
d_loss_d_z2 = d_loss_d_a2 * d_a2_d_z2  # (1,)
print(f"d_loss_d_z2 = {d_loss_d_z2}")

# No.二layerParameter梯degree # Second layer parameter gradients
d_z2_d_W2 = a1  # (3,), because z2 = W2@a1 + b2 # (3,), because z2 = W2@a1 + b2
d_z2_d_b2 = 1.0  # (1,)
d_z2_d_a1 = W2[0]  # (3,)

# Compute parameter gradients # Compute parameter gradients
d_loss_d_W2 = d_loss_d_z2.reshape(-1, 1) @ d_z2_d_W2.reshape(1, -1)  # (1, 3)
d_loss_d_b2 = d_loss_d_z2 * d_z2_d_b2  # (1,)
print(f"d_loss_d_W2 = {d_loss_d_W2}")
print(f"d_loss_d_b2 = {d_loss_d_b2}")

# Gradient passed to the hidden layer # Gradient passed to the hidden layer
d_loss_d_a1 = d_loss_d_z2 @ d_z2_d_a1  # (3,)
print(f"d_loss_d_a1 = {d_loss_d_a1}")

# Hidden layer gradient # Hidden layer gradient
d_a1_d_z1 = relu_derivative(z1)  # (3,)
d_loss_d_z1 = d_loss_d_a1 * d_a1_d_z1  # (3,)
print(f"d_loss_d_z1 = {d_loss_d_z1}")

# First layer parameter gradients # First layer parameter gradients
d_z1_d_W1 = x  # (2,)
d_z1_d_b1 = 1.0  # (3,)

# Compute parameter gradients # Compute parameter gradients
d_loss_d_W1 = d_loss_d_z1.reshape(-1, 1) @ d_z1_d_W1.reshape(1, -1)  # (3, 2)
d_loss_d_b1 = d_loss_d_z1 * d_z1_d_b1  # (3,)
print(f"d_loss_d_W1 = {d_loss_d_W1}")
print(f"d_loss_d_b1 = {d_loss_d_b1}")

# ============================================
# Parameter update # Parameter update
# ============================================
print("\n=== Parameter update === === Parameter update ===)
learning_rate = 0.01

W1_new = W1 - learning_rate * d_loss_d_W1
b1_new = b1 - learning_rate * d_loss_d_b1
W2_new = W2 - learning_rate * d_loss_d_W2
b2_new = b2 - learning_rate * d_loss_d_b2

print("Parameters have been updated." "Parameters have been updated.")

ThisExamplesshows反toward传播ofCompleteProcess. This example demonstrates the complete process of backpropagation.

The core rule is: The core rules are:

  • forLine性layer y = Wx + b:dL/dW = (dL/dy)·xᵀ,dL/db = dL/dy For a linear layer y = Wx + b: dL/dW = (dL/dy)·xᵀ, dL/db = dL/dy
  • foractivateFunction y = f(z):dL/dz = dL/dy ⊙ f'(z)(⊙ Yes逐elementMultiplication) For an activation function y = f(z): dL/dz = dL/dy ⊙ f'(z) (⊙ is element-wise multiplication)
  • 梯degreefromafter往before传,everyonelayeralluse链formula法Then"积累"梯degree Gradients propagate from back to front, and each layer uses the chain rule to "accumulate" gradients.

OKMessageYes:presentgenerationDeep LearningFramework(such as PyTorch、TensorFlow)willAutomaticisyouCompute反toward传播。you只requiresdefinitionbeforetoward传播,FrameworkKnow how to useAutomaticmicroDivide(autograd)搞decideAll梯degreeCalculate. The good news is: modern deep learning frameworks (such as PyTorch, TensorFlow) automatically compute backpropagation for you. You only need to define the forward pass, and the framework will handle all gradient computation with automatic differentiation (autograd).


Gradient Descent Optimizers

Compute出梯degreeAfter,I们requiresuseExcellenttransformdevicecomeUpdateparameters. After computing the gradients, we need an optimizer to update the parameters.

SimplestofExcellenttransformdeviceYesRandom梯degreeBottom降(SGD), but还Yes很many更AdvancedofExcellenttransformdevice. The simplest optimizer is stochastic gradient descent (SGD), but there are many more advanced optimizers.

SGD: Stochastic Gradient Descent

Formula: θ = θ - η·∇L(θ) Formula: θ = θ - η·∇L(θ)

where η is the learning rate. where η is the learning rate.

Example

# ============================================
# SGD optimizer # SGD optimizer
# ============================================

import numpy as np

def f(x):
    """I们want最smalltransformofFunction:f(x) = x²""" """The function we want to minimize: f(x) = x²"""
    return x ** 2

def f_grad(x):
    """梯degree:f'(x) = 2x""" """Gradient: f'(x) = 2x"""
    return 2 * x

# SGD optimization # SGD optimization
x = 10.0  # Initial value # Initial value
learning_rate = 0.1  # Learning rate # Learning rate
steps = 20

print(f"初start x = {x}, f(x) = {f(x)}" "Initial x = {x}, f(x) = {f(x)}")
print("-" * 40)

for step in range(steps):
    grad = f_grad(x)
    x = x - learning_rate * grad  # SGD update # SGD update
    print(f"Step {step+1}: x = {x:.6f}, f(x) = {f(x):.6f}")

print("-" * 40)
print(f"最终 x = {x:.6f}, f(x) = {f(x):.6f}" "Final x = {x:.6f}, f(x) = {f(x):.6f}")
print("manage论最Excellent:x = 0, f(x) = 0" "Theoretical optimum: x = 0, f(x) = 0")

SGD SimpleButYesDisadvantages: SGD is simple but has drawbacks:

  • Uses the same learning rate for all parameters
  • Convergence can be slow
  • Easily gets stuck in local optima or saddle points.
  • Sensitive to noise

Momentum: Acceleration with Momentum

Momentum 引入already"Speed"ofConcepts——Let梯degreeUpdate带Yes惯性。 Momentum introduces the concept of "velocity" — giving gradient updates inertia.

Formula: Formula:

v = β·v + (1-β)·∇L(θ)
θ = θ - η·v

β is usually set to 0.9. β is usually set to 0.9.

Example

# ============================================
# Momentum optimizer # Momentum optimizer
# ============================================

import numpy as np

def f(x):
    """Objective function""" """Objective function"""
    return x ** 2

def f_grad(x):
    """Gradient""" """Gradient"""
    return 2 * x

def sgd_optimize(x_init, learning_rate, steps):
    """Pure SGD""" """Pure SGD"""
    x = x_init
    history = [x]
    for _ in range(steps):
        grad = f_grad(x)
        x = x - learning_rate * grad
        history.append(x)
    return np.array(history)

def momentum_optimize(x_init, learning_rate, beta, steps):
    """Momentum"""
    x = x_init
    v = 0.0  # Initial velocity # Initial velocity
    history = [x]
    for _ in range(steps):
        grad = f_grad(x)
        v = beta * v + (1 - beta) * grad
        x = x - learning_rate * v
        history.append(x)
    return np.array(history)

# Comparison # Comparison
x_init = 10.0
learning_rate = 0.1
beta = 0.9
steps = 20

sgd_history = sgd_optimize(x_init, learning_rate, steps)
momentum_history = momentum_optimize(x_init, learning_rate, beta, steps)

print("Step |    SGD    |  Momentum")
print("-" * 35)
for i in range(steps + 1):
    print(f"{i:4} | {sgd_history[i]:9.6f} | {momentum_history[i]:9.6f}")

print("\nObservation: Momentum converges faster!)
print(Because it accumulates past gradients, giving it inertia.)

Adam: Adaptive Learning Rate

Adam (Adaptive Moment Estimation) is currently one of the most commonly used optimizers.

It combines the ideas of Momentum (first moment) and RMSProp (second moment) to maintain an adaptive learning rate for each parameter.

Formula:

m = β₁·m + (1-β₁)·∇L(θ)
v = β₂·v + (1-β₂)·(∇L(θ))²
m̂ = m / (1-β₁ᵗ)
v̂ = v / (1-β₂ᵗ)
θ = θ - η·m̂ / (√v̂ + ε)

Default hyperparameters: β₁=0.9, β₂=0.999, ε=1e-8, η=1e-3.

Example

# ============================================
# Adam optimizer
# ============================================

import numpy as np

def adam_optimize(x_init, learning_rate, beta1, beta2, epsilon, steps):
    """Adam optimization"""
    x = x_init
    m = 0.0  # First moment
    v = 0.0  # Second moment
    history = [x]

    for t in range(1, steps + 1):
        grad = 2 * x  # Gradient of f(x) = x²
        # Update first and second moments
        m = beta1 * m + (1 - beta1) * grad
        v = beta2 * v + (1 - beta2) * (grad ** 2)
        # Bias correction
        m_hat = m / (1 - beta1 ** t)
        v_hat = v / (1 - beta2 ** t)
        # Parameter update
        x = x - learning_rate * m_hat / (np.sqrt(v_hat) + epsilon)
        history.append(x)

    return np.array(history)

# Parameters
x_init = 10.0
learning_rate = 0.5
beta1 = 0.9
beta2 = 0.999
epsilon = 1e-8
steps = 20

# Optimization
history = adam_optimize(x_init, learning_rate, beta1, beta2, epsilon, steps)

print("Adam optimization process:")
print("Step |     x     |   f(x)")
print("-" * 35)
for i in range(steps + 1):
    print(f"{i:4} | {history[i]:9.6f} | {history[i]**2:9.6f}")

AdamW: Weight Decay Correction

AdamW is an improved version of Adam — it separates weight decay from the gradient.

This is useful for regularization.

Optimizer comparison:

OptimizerFormulaAdvantagesDisadvantagesApplicable scenarios
SGDθ = θ - η·∇LSimple, stableSlow convergence, sensitive to learning rateLarge datasets, rich experience in hyperparameter tuning
SGD + Momentumv = βv + (1-β)∇L
θ = θ - ηv
Accelerates convergence, reduces oscillationRequires tuning βGenerally better than plain SGD
AdamAdaptive learning rateFast convergence, requires less hyperparameter tuningMay generalize worse than SGDMost scenarios (default choice)
AdamWAdam + weight decayBetter regularizationRequires tuning weight decay coefficientTransformers, large models

Optimizer selection suggestion: Try Adam first, with the learning rate set to 1e-3 (3e-4 is a safer choice). If results are unsatisfactory, try tuning hyperparameters or switching to another optimizer.


Regularization Techniques

The goal of regularization is to prevent overfitting — helping the model perform well on unseen data.

L1/L2 Regularization

L1 and L2 regularization are the most basic regularization methods.

They add a penalty on parameters to the loss function:

L_total = L + λ·L_reg

L1 regularization (Lasso): L_reg = |w|

L2 regularization (Ridge): L_reg = w²

Example

# ============================================
# L1/L2 regularization
# ============================================

import numpy as np

def l1_regularization(weights, lambda_l1):
    """L1 regularization: returns the regularization term and gradient"""
    reg = lambda_l1 * np.sum(np.abs(weights))
    grad = lambda_l1 * np.sign(weights)
    return reg, grad

def l2_regularization(weights, lambda_l2):
    """L2 regularization: returns the regularization term and gradient"""
    reg = 0.5 * lambda_l2 * np.sum(weights ** 2)
    grad = lambda_l2 * weights
    return reg, grad

# Example
weights = np.array([1.0, -2.0, 3.0, 0.5])
print(f"Weights: {weights}")

lambda_l1 = 0.1
lambda_l2 = 0.1

l1_reg, l1_grad = l1_regularization(weights, lambda_l1)
print(f"\nL1 regularization term: {l1_reg:.4f}")
print(f"L1 gradient: {l1_grad}")

l2_reg, l2_grad = l2_regularization(weights, lambda_l2)
print(f"\nL2 regularization term: {l2_reg:.4f}")
print(f"L2 gradient: {l2_grad}")

print("\nL1 vs L2:")
print("- L1: tends to produce sparse solutions (many parameters become 0), can be used for feature selection")
print("- L2: tends to make all parameters very small, but not zero")
print("- In deep learning, L2 is more commonly used (or weight decay via AdamW)")

Dropout

Dropout is a simple but effective regularization technique: randomly "turning off" a portion of neurons during training.

This forces the model to learn more robust features — it cannot rely on any specific neuron.

Example

# ============================================
# Dropout
# ============================================

import numpy as np

def dropout_forward(x, dropout_rate, training=True):
    """Dropout forward pass"""
    if not training:
        # No dropout during testing, return directly
        return x

    # Generate mask: randomly set a portion of neurons to 0
    mask = (np.random.rand(*x.shape) > dropout_rate).astype(float)
    # Scaling: no scaling needed at test time, so divide by (1 - dropout_rate) during training
    scale = 1.0 / (1.0 - dropout_rate)
    # Apply dropout
    out = x * mask * scale

    return out, mask

# Test
x = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
dropout_rate = 0.4

print(f"Input x: {x}")
print(f"Dropout rate: {dropout_rate}")

# Run multiple times to observe randomness
print("\nMultiple Dropout results:")
for i in range(3):
    out, mask = dropout_forward(x, dropout_rate, training=True)
    print(f" Round {i+1}: mask={mask}, out={out}")

# Test mode
out_test = dropout_forward(x, dropout_rate, training=False)
print(f"\nTest mode (no Dropout): {out_test}")

Data Augmentation

Data augmentation is one of the most effective regularization methods—it creates more training samples by transforming existing data.

For images, common data augmentation techniques include: flipping, rotation, scaling, cropping, color jittering, etc.

For text, common data augmentation techniques include: synonym replacement, random insertion, random deletion, back-translation, etc.

Summary of regularization techniques:

MethodPrincipleAdvantagesDisadvantages
L2 RegularizationPenalizes large weightsSimple and stableLimited effect
DropoutRandomly drops neuronsSimple and effectiveSlows down training
Data AugmentationIncreases diversity of training dataSignificant effectRequires domain knowledge
Early StoppingStop when validation set stops improvingSimple and effectiveRequires monitoring the validation set
Batch NormalizationNormalize layer inputsAccelerates training, regularizesAdds computational overhead

Batch Normalization vs Layer Normalization

Normalization techniques can speed up training and make optimization more stable.

Internal Covariate Shift

When training deep networks, the input distribution of each layer changes as the parameters of the previous layer change. This is "internal covariate shift" (Internal Covariate Shift).

This slows down training because each layer has to constantly adapt to new distributions.

The goal of normalization techniques is to keep the input distribution of each layer stable.

Batch Normalization

Batch Normalization (BatchNorm) normalizes along the batch dimension.

Formula:

μ = 1/N · sum(xᵢ)
σ² = 1/N · sum((xᵢ - μ)²)
x̂ᵢ = (xᵢ - μ) / √(σ² + ε)
yᵢ = γ·x̂ᵢ + β

where γ and β are learnable parameters.

Example

# ============================================
# Batch Normalization
# ============================================

import numpy as np

def batch_norm(x, gamma, beta, epsilon=1e-5, training=True, running_mean=None, running_var=None, momentum=0.1):
    """Batch Normalization
x: input, shape (batch_size, features)
gamma: scaling parameter
beta: offset parameter
    """

    if training:
        # Training mode: use statistics of the current batch
        batch_mean = np.mean(x, axis=0)
        batch_var = np.var(x, axis=0)

        # Update running statistics (for use at test time)
        if running_mean is not None and running_var is not None:
            running_mean[:] = momentum * running_mean + (1 - momentum) * batch_mean
            running_var[:] = momentum * running_var + (1 - momentum) * batch_var

        # Normalize
        x_normalized = (x - batch_mean) / np.sqrt(batch_var + epsilon)
    else:
        # Test mode: use running statistics
        x_normalized = (x - running_mean) / np.sqrt(running_var + epsilon)

    # Scale and shift
    out = gamma * x_normalized + beta
    return out

# Test
batch_size = 4
features = 3

# One batch of data
x = np.array([
    [1.0, 2.0, 3.0],
    [2.0, 3.0, 4.0],
    [3.0, 4.0, 5.0],
    [4.0, 5.0, 6.0],
])
print(f"Input x:\n{x}")

# Initial parameters
gamma = np.ones(features)  # Initially no scaling
beta = np.zeros(features)  # Initially no offset
print(f"\ngamma: {gamma}")
print(f"beta: {beta}")

# Apply batch normalization
out = batch_norm(x, gamma, beta, training=True)
print(f"\nAfter batch normalization:\n{out}")

# Verify: mean of each feature should be close to 0, variance close to 1
print(f"\nNormalized mean: {np.mean(out, axis=0)}")
print(f"Normalized variance: {np.var(out, axis=0)}")

Advantages of batch normalization:

  • Speeds up training
  • Allows larger learning rates
  • Reduces sensitivity to initialization
  • Has a slight regularization effect

Disadvantages of batch normalization:

  • Depends on batch size—works poorly with too-small batches
  • Behavior differs between training and testing
  • Not suitable for sequence models (e.g., RNN, Transformer)

Layer Normalization

Layer Normalization (LayerNorm) normalizes along the feature dimension instead of the batch dimension.

This makes it independent of batch size and very suitable for sequence models.

Example

# ============================================
# Layer Normalization
# ============================================

import numpy as np

def layer_norm(x, gamma, beta, epsilon=1e-5):
    """Layer Normalization
x: input, shape (batch_size, features) or (batch_size, seq_len, features)
    """

    # Compute mean and variance over the last dimension (feature dimension)
    mean = np.mean(x, axis=-1, keepdims=True)
    var = np.var(x, axis=-1, keepdims=True)

    # Normalize
    x_normalized = (x - mean) / np.sqrt(var + epsilon)

    # Scale and shift
    out = gamma * x_normalized + beta
    return out

# Test
batch_size = 2
seq_len = 3
features = 4

# Simulate input in Transformer: (batch_size, seq_len, features)
x = np.array([
    [[1.0, 2.0, 3.0, 4.0],
     [2.0, 3.0, 4.0, 5.0],
     [3.0, 4.0, 5.0, 6.0]],
    [[4.0, 5.0, 6.0, 7.0],
     [5.0, 6.0, 7.0, 8.0],
     [6.0, 7.0, 8.0, 9.0]],
])
print(f"Input x shape: {x.shape}")

# Initialize parameters
gamma = np.ones(features)
beta = np.zeros(features)

# Apply layer normalization
out = layer_norm(x, gamma, beta)
print(f"Shape after layer normalization: {out.shape}")

# Verify: mean at each position is close to 0, variance close to 1
mean_check = np.mean(out, axis=-1)
var_check = np.var(out, axis=-1)
print(f"\nMean at each position:\n{mean_check}")
print(f"Variance at each position:\n{var_check}")

Comparison of the two normalization methods:

MethodNormalization dimensionDepends on batchApplicable scenariosRepresentative models
BatchNormbatch dimensionYesCNN, fixed batch sizeResNet、VGG
LayerNormfeature dimensionnoTransformer, RNN, variable-length sequencesGPT、BERT、LLaMA

For Transformer and large language models, LayerNorm is the standard choice. It computes statistics independently at each position, does not depend on batch size, and does not need to maintain running statistics.


Learning Rate Scheduling

Learning rate is one of the most important hyperparameters — too large and it won't converge, too small and convergence is too slow.

Learning rate scheduling is dynamically adjusting the learning rate during training.

Warmup

In the early stage of training, use a small learning rate to "warm up" and let the model stabilize, then increase to the target learning rate.

This is especially important for Transformer.

Cosine Annealing

Cosine annealing gradually decreases the learning rate following the shape of the cosine function.

Example

# ============================================
# Learning rate scheduling
# ============================================

import numpy as np

def warmup_schedule(step, warmup_steps, base_lr):
    """Linear warmup schedule"""
    if step < warmup_steps:
        return base_lr * (step + 1) / warmup_steps
    return base_lr

def cosine_annealing_schedule(step, total_steps, base_lr, min_lr=0.0):
    """Cosine annealing schedule"""
    return min_lr + 0.5 * (base_lr - min_lr) * (1 + np.cos(np.pi * step / total_steps))

def warmup_cosine_schedule(step, warmup_steps, total_steps, base_lr, min_lr=0.0):
    """Warmup + cosine annealing"""
    if step < warmup_steps:
        return base_lr * (step + 1) / warmup_steps
    return min_lr + 0.5 * (base_lr - min_lr) * (1 + np.cos(np.pi * (step - warmup_steps) / (total_steps - warmup_steps)))

# Parameters
total_steps = 100
warmup_steps = 10
base_lr = 1e-3
min_lr = 1e-5

print("Different learning rate schedules:")
print("Step |  Warmup  |  Cosine  |  Warmup+Cosine")
print("-" * 55)

for step in range(0, total_steps + 1, 10):
    lr_warmup = warmup_schedule(step, warmup_steps, base_lr)
    lr_cosine = cosine_annealing_schedule(step, total_steps, base_lr, min_lr)
    lr_combined = warmup_cosine_schedule(step, warmup_steps, total_steps, base_lr, min_lr)
    print(f"{step:4} | {lr_warmup:.6f} | {lr_cosine:.6f} | {lr_combined:.6f}")

The Impact of Learning Rate on Training

The impact of learning rate:

  • Too large: training unstable, may diverge
  • Too small: slow convergence, may get stuck in local optima
  • Appropriate: fast and stable convergence

A common strategy is the "learning rate finder": first train a few steps with learning rates increasing from small to large, and see at which learning rate the loss decreases fastest.

Summary of learning rate scheduling methods:

Scheduling methodCharacteristicsApplicable scenarios
Fixed learning rateSimplestSmall models, simple tasks
Step DecayDecrease every few stepsCNN, etc.
Cosine AnnealingSmooth decreaseMost scenarios
Warmup + CosineIncrease first then decreaseTransformer, large models (recommended)
Reduce on PlateauDecrease when validation set does not improveRequires monitoring the validation set

Hands-On: Implementing an MLP from Scratch with PyTorch

Now let's integrate all the previous knowledge points, use PyTorch to implement a complete multi-layer perceptron (MLP), and train it on a simple classification task.

Example

# ============================================
# Complete MLP training pipeline implemented in PyTorch
# ============================================

import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import Dataset, DataLoader
import numpy as np

# ============================================
# 1. Prepare data
# ============================================

class SimpleDataset(Dataset):
    """A simple binary classification dataset"""

    def __init__(self, n_samples=1000, n_features=10):
        # Generate synthetic data
        np.random.seed(42)
        self.X = np.random.randn(n_samples, n_features).astype(np.float32)
        # Simple classification rule: y = 1 if sum(x[:5]) > 0 else 0
        self.y = (np.sum(self.X[:, :5], axis=1) > 0).astype(np.int64)

    def __len__(self):
        return len(self.X)

    def __getitem__(self, idx):
        return torch.tensor(self.X[idx]), torch.tensor(self.y[idx])

# Create dataset and data loader
train_dataset = SimpleDataset(n_samples=800)
val_dataset = SimpleDataset(n_samples=200)

train_loader = DataLoader(train_dataset, batch_size=32, shuffle=True)
val_loader = DataLoader(val_dataset, batch_size=32, shuffle=False)

print(f"Training set size: {len(train_dataset)}")
print(f"Validation set size: {len(val_dataset)}")

# ============================================
# 2. Define the model
# ============================================

class MLP(nn.Module):
    """Multi-layer perceptron"""

    def __init__(self, input_dim, hidden_dims, output_dim, dropout_rate=0.1):
        super().__init__()

        layers = []
        prev_dim = input_dim

        # Hidden layer
        for hidden_dim in hidden_dims:
            layers.append(nn.Linear(prev_dim, hidden_dim))
            layers.append(nn.ReLU())
            layers.append(nn.LayerNorm(hidden_dim))  # Layer normalization
            layers.append(nn.Dropout(dropout_rate))  # Dropout
            prev_dim = hidden_dim

        # Output layer
        layers.append(nn.Linear(prev_dim, output_dim))

        self.model = nn.Sequential(*layers)

    def forward(self, x):
        return self.model(x)

# Create model
model = MLP(
    input_dim=10,
    hidden_dims=[64, 32],
    output_dim=2,
    dropout_rate=0.1
)
print("\nModel structure:")
print(model)

# ============================================
# 3. Define loss function and optimizer
# ============================================

criterion = nn.CrossEntropyLoss()  # Cross-entropy loss
optimizer = optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-4)  # AdamW optimizer

# Learning rate scheduler: warmup + cosine annealing
total_epochs = 20
warmup_epochs = 2
total_steps = len(train_loader) * total_epochs
warmup_steps = len(train_loader) * warmup_epochs

def lr_lambda(step):
    if step < warmup_steps:
        return (step + 1) / warmup_steps
    else:
        progress = (step - warmup_steps) / (total_steps - warmup_steps)
        return 0.5 * (1 + np.cos(np.pi * progress))

scheduler = optim.lr_scheduler.LambdaLR(optimizer, lr_lambda=lr_lambda)

print("\nTraining configuration:")
print(f"  Epochs: {total_epochs}")
print(f"  Warmup epochs: {warmup_epochs}")
print(f"  Total steps: {total_steps}")
print(f"  Warmup steps: {warmup_steps}")

# ============================================
# 4. Training loop
# ============================================

print("\nStart training...")
print("-" * 50)

global_step = 0
best_val_acc = 0.0

for epoch in range(total_epochs):
    # Training phase
    model.train()
    train_loss = 0.0
    train_correct = 0
    train_total = 0

    for batch_X, batch_y in train_loader:
        # Forward pass
        logits = model(batch_X)
        loss = criterion(logits, batch_y)

        # Backward pass
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()
        scheduler.step()

        # Statistics
        train_loss += loss.item()
        predicted = torch.argmax(logits, dim=1)
        train_total += batch_y.size(0)
        train_correct += (predicted == batch_y).sum().item()

        global_step += 1

    train_loss /= len(train_loader)
    train_acc = train_correct / train_total

    # Validation phase
    model.eval()
    val_loss = 0.0
    val_correct = 0
    val_total = 0

    with torch.no_grad():
        for batch_X, batch_y in val_loader:
            logits = model(batch_X)
            loss = criterion(logits, batch_y)

            val_loss += loss.item()
            predicted = torch.argmax(logits, dim=1)
            val_total += batch_y.size(0)
            val_correct += (predicted == batch_y).sum().item()

    val_loss /= len(val_loader)
    val_acc = val_correct / val_total

    # Save the best model
    if val_acc > best_val_acc:
        best_val_acc = val_acc
        # You can save model checkpoint here

    # Print log
    current_lr = optimizer.param_groups[0]['lr']
    print(f"Epoch {epoch+1:2d} | lr: {current_lr:.6f} | "
          f"Train Loss: {train_loss:.4f} | Train Acc: {train_acc:.4f} | "
          f"Val Loss: {val_loss:.4f} | Val Acc: {val_acc:.4f}")

print("-" * 50)
print(f"Training complete! Best validation accuracy: {best_val_acc:.4f}")

# ============================================
# 5. Inference example
# ============================================

print("\nInference example:")
model.eval()

# A few test samples
test_samples = [
    np.array([1.0, 1.0, 1.0, 1.0, 1.0, 0.0, 0.0, 0.0, 0.0, 0.0], dtype=np.float32),  # Should be 1
    np.array([-1.0, -1.0, -1.0, -1.0, -1.0, 0.0, 0.0, 0.0, 0.0, 0.0], dtype=np.float32),  # Should be 0
]

for i, sample in enumerate(test_samples):
    with torch.no_grad():
        input_tensor = torch.tensor(sample).unsqueeze(0)
        logits = model(input_tensor)
        probs = torch.softmax(logits, dim=1)
        predicted_class = torch.argmax(probs, dim=1).item()

    print(f" Sample {i+1}: {sample[:5]}...")
    print(f" Predicted class: {predicted_class}")
    print(f" Probabilities: class 0={probs[0,0]:.4f}, class 1={probs[0,1]:.4f}")

This example contains all standard components of deep learning training:

  • Data processing: Dataset, DataLoader
  • Model definition: MLP with LayerNorm and Dropout
  • Loss function: CrossEntropyLoss
  • Optimizer: AdamW
  • Learning rate scheduler: Warmup + Cosine
  • Training loop: training phase + validation phase
  • Model saving and inference

This workflow can be directly extended to more complex tasks and models.

Other extensions