Common Optimizers -- SGD, Momentum, and Adam
Basic gradient descent has multiple improved versions in practice. This chapter compares the three most commonly used optimizers.
Concept Analysis
SGD
Randomly sample a mini-batch each time to compute the gradient
\( \theta = \theta - \eta g \)
Simple, but large oscillations
Momentum
Accumulates historical directions, like rolling a snowball
\( v = \beta v + \eta g \)
\( \theta = \theta - v \)
Speeds up convergence, reduces oscillations
Adam
Momentum + adaptive learning rate
Each parameter has its own learning rate
Default first choice, works well for most tasks
Applicable Scenarios for Each Optimizer
| Optimizer | Use Case |
|---|---|
| SGD + Momentum | CV tasks, ResNet training |
| Adam | Default choice for Transformer/BERT/GPT |
| AdamW | De facto standard for large model training |
Everyday Examples
SGD = Only looks at a small section of the road each time to decide the next direction (efficient but a bit shaky).
Momentum = Remembers the previous direction, like a snowball accelerating downhill.
Adam = Uses different step sizes for different directions (parameters), small steps on steep slopes, large steps on flat areas.
Python Hands-On Practice
Example
def loss(w, b):
return (w-3)**2 + 2*(b+2)**2 # Optimal point (3, -2)
# SGD
def run_SGD(lr=0.1, steps=30):
w, b = 0.0, 0.0
losses = []
for _ in range(steps):
w -= lr * 2*(w-3); b -= lr * 4*(b+2)
losses.append(loss(w, b))
return losses
# Momentum
def run_Momentum(lr=0.1, beta=0.9, steps=30):
w, b, vw, vb = 0.0, 0.0, 0.0, 0.0
losses = []
for _ in range(steps):
vw = beta*vw + lr*2*(w-3); vb = beta*vb + lr*4*(b+2)
w -= vw; b -= vb
losses.append(loss(w, b))
return losses
print("=== EXAMPLE Optimizer Convergence Comparison ===")
for name, fn in [('SGD', run_SGD), ('Momentum', run_Momentum)]:
l = fn()
print(f"{name:10s}: initial loss={l[0]:.1f}, final loss={l[-1]:.4f}")
Run output:
=== EXAMPLE 优化器收敛对比 === SGD : 初始loss=25.0, 最终loss=1.9234 Momentum : 初始loss=25.0, 最终loss=0.0021
Convergence Curve Comparison
Other Extensions