Gradient Descent -- Descending along the steepest direction

This chapter implements gradient descent from scratch and uses it to fit linear regression.This is the mathematical prototype of the entire AI training loop.


Concept Analysis

Core Formula

\[ \theta_{t+1} = \theta_t - \eta \nabla J(\theta_t) \]

The Four-Step Training Loop


Forward Propagation
Predict Output

Compute Loss
Measure the Gap

Compute Gradient
Backpropagation

Update Parameters
θ -= η·∇J

Impact of Learning Rate

Learning RateEffect
Too SmallConverges extremely slowly, requiring many steps
ModerateConverges smoothly and quickly
Too LargeOscillates or even diverges

Real-life Example

Descending the Mountain on a Foggy Day

On a foggy day where you can't see your hand in front of your face, you need to descend the mountain. The strategy for each step:

Feel the steepest direction under your feet → Take a small step in the steepest downhill direction → Stop → Feel again → Take another small step → Repeat.

This is gradient descent: feeling = computing the gradient, taking a small step = updating parameters, stopping = next iteration.


Python Hands-On Practice

Example

import numpy as np

# Generate data y = 3x + 2 + noise
np.random.seed(42)
X = np.linspace(0, 10, 100)
y = 3 * X + 2 + np.random.normal(0, 2, 100)

# Implement gradient descent from scratch
w, b = 0.0, 0.0
lr, n_iter = 0.01, 200
losses = []

for i in range(n_iter):
    y_pred = w * X + b
    loss = np.mean((y_pred - y) ** 2)
    losses.append(loss)
    # Gradient
    dw = 2 * np.mean((y_pred - y) * X)
    db = 2 * np.mean(y_pred - y)
    # Update
    w -= lr * dw
    b -= lr * db

print(f"EXAMPLE gradient descent result:")
print(f"True: w=3.0, b=2.0")
print(f"Fitted: w={w:.4f}, b={b:.4f}")
print(fInitial loss: {losses[0]:.2f} → Final loss: {losses[-1]:.2f})
EXAMPLE 梯度下降结果:
真实: w=3.0, b=2.0
拟合: w=2.9874, b=2.2835
初始损失: 79.69 → 最终损失: 3.91

Loss Descent Curve


Application Scenarios in AI

The Mathematical Prototype of the PyTorch Training Loop

All deep learning training loops follow the same pattern, which is gradient descent:

① optimizer.zero_grad() → zero the gradient cache; ② loss = model(x) → forward propagation; ③ loss.backward() → backpropagation to compute gradients; ④ optimizer.step() → perform parameter update θ -= η·∇J.

Once you understand this four-step loop, you understand the core training logic of all deep learning frameworks.

Learning Rate Is the Most Important Hyperparameter

Learning rate too large → loss oscillates or even diverges; learning rate too small → convergence is extremely slow and may get stuck in local optima. Learning rate scheduling (Step, Cosine Annealing, Warmup) is an indispensable technique for training large models — GPT-3 training used a linear warmup + cosine decay strategy.

SGD Batch Size Trade-off

Full-batch GD: uses all data to compute exact gradients, slow but stable; SGD (batch_size=1): looks at only one sample at a time, fast but gradient estimates are noisy; Mini-batch SGD: a compromise (batch_size=32/64/128), with moderate noise, which actually helps escape local optima — this is the standard practice in real-world training.

Limitations of Gradient Descent

Gradient descent finds points where the gradient is zero — these may be global optima, local optima, or saddle points. In high-dimensional spaces, saddle points are far more common than local optima: a point is a minimum in some directions and a maximum in others. Adaptive optimizers such as Adam can escape saddle points more effectively.


Other extensions