Common Optimizers -- SGD, Momentum, and Adam

Basic gradient descent has multiple improved versions in practice. This chapter compares the three most commonly used optimizers.


Concept Analysis

SGD

Randomly sample a mini-batch each time to compute the gradient

\( \theta = \theta - \eta g \)

Simple, but large oscillations

Momentum

Accumulates historical directions, like rolling a snowball

\( v = \beta v + \eta g \)

\( \theta = \theta - v \)

Speeds up convergence, reduces oscillations

Adam

Momentum + adaptive learning rate

Each parameter has its own learning rate

Default first choice, works well for most tasks

Applicable Scenarios for Each Optimizer

OptimizerUse Case
SGD + MomentumCV tasks, ResNet training
AdamDefault choice for Transformer/BERT/GPT
AdamWDe facto standard for large model training

Everyday Examples

SGD = Only looks at a small section of the road each time to decide the next direction (efficient but a bit shaky).

Momentum = Remembers the previous direction, like a snowball accelerating downhill.

Adam = Uses different step sizes for different directions (parameters), small steps on steep slopes, large steps on flat areas.


Python Hands-On Practice

Example

import numpy as np

def loss(w, b):
    return (w-3)**2 + 2*(b+2)**2  # Optimal point (3, -2)

# SGD
def run_SGD(lr=0.1, steps=30):
    w, b = 0.0, 0.0
    losses = []
    for _ in range(steps):
        w -= lr * 2*(w-3); b -= lr * 4*(b+2)
        losses.append(loss(w, b))
    return losses

# Momentum
def run_Momentum(lr=0.1, beta=0.9, steps=30):
    w, b, vw, vb = 0.0, 0.0, 0.0, 0.0
    losses = []
    for _ in range(steps):
        vw = beta*vw + lr*2*(w-3); vb = beta*vb + lr*4*(b+2)
        w -= vw; b -= vb
        losses.append(loss(w, b))
    return losses

print("=== EXAMPLE Optimizer Convergence Comparison ===")
for name, fn in [('SGD', run_SGD), ('Momentum', run_Momentum)]:
    l = fn()
    print(f"{name:10s}: initial loss={l[0]:.1f}, final loss={l[-1]:.4f}")

Run output:

=== EXAMPLE 优化器收敛对比 ===
SGD       : 初始loss=25.0, 最终loss=1.9234
Momentum  : 初始loss=25.0, 最终loss=0.0021

Convergence Curve Comparison


Other Extensions