Gradient -- points in the direction of the fastest increase of the function.

The gradient points in the direction in which the function value increases fastest, so gradient descent goes in the negative gradient direction.

Gradient is the core concept of the entire AI training process — every update step taken by the optimizer relies on the gradient.


Concept Analysis

Gradient = a vector composed of partial derivatives

\[ \nabla f(x, y) = \left( \frac{\partial f}{\partial x}, \frac{\partial f}{\partial y} \right) \]

Gradient Direction

Points in the direction of steepest ascent

Want to go uphill → follow the gradient

Negative gradient direction

Points in the direction of steepest descent

Want to go downhill → follow the negative gradient

Gradient Magnitude

The steepness of that direction

Gradient is zero → extremum or saddle point

The negative sign in the update formula \( \theta = \theta - \eta \nabla J \) means "move in the opposite direction (downhill)."

If you forget the negative sign, it becomes gradient ascent, and the loss will get larger and larger.


Everyday Examples

Descending the mountain on a foggy day

You are standing on a mountain shrouded in fog. To go down as quickly as possible, you feel which direction under your feet is steepest, then take a step in the steepest downhill direction.

Repeat continuously: feel → step → feel → step. This is the physical intuition of gradient descent.

The gradient tells you the "steepest direction," and the negative sign tells you to "go down."


Mathematical Definition

\[ \nabla f(\mathbf{x}) = \begin{bmatrix} \frac{\partial f}{\partial x_1} \\ \vdots \\ \frac{\partial f}{\partial x_n} \end{bmatrix} \]

direction导number沿 u direction:\( \nabla f \cdot \mathbf{u} = \|\nabla f\| \cos\theta \)。When u and梯degreesametowardwhenmaximum。


Hands-on Python practice

Example

import numpy as np

def f(x, y):
    return x**2 + y**2  # Bowl-shaped function

def grad(x, y):
    return np.array([2*x, 2*y])

pt = np.array([1.5, 1.0])
g = grad(*pt)
print(fAt (1.5, 1.0): gradient = {g})
print(fGradient norm = {np.linalg.norm(g):.2f})
print(fThe gradient points away from the origin (ascent), and the negative gradient points toward the origin (descent))

Interactive 3D visualization of the loss surface

Below we use Plotly.js to draw the 3D surface of \( f(x,y)=x^2+y^2 \) to intuitively feel the gradient direction:


Application scenarios in AI

Core input for gradient descent training

Every parameter update \( \theta = \theta - \eta \nabla_\theta J \) depends on the gradient. The gradient tells each parameter "increase or decrease, and by how much." In PyTorch, loss.backward() computes the gradient, and optimizer.step() uses the gradient to update the parameters.

Vanishing gradients and exploding gradients

In deep networks, during backpropagation the gradient may become increasingly small (vanishing, causing shallow layers to learn nothing) or increasingly large (exploding, causing parameter updates to be too large and diverge). ResNet's residual connections (identity shortcut) are designed to alleviate gradient vanishing—giving the gradient a "highway" that goes straight to the shallow layers.

Gradient Clipping

When training RNNs and large language models, gradient clipping is commonly applied: if the L2 norm of the gradient exceeds a threshold, it is scaled down to the threshold size. This is a practical technique to prevent gradient explosion; the standard practice in GPT training is clip_grad_norm_=1.0.

Practical significance of visualizing gradients

Tools like TensorBoard and W&B can visualize the distribution of gradients in each layer. If the gradient of a layer is almost zero → gradient vanishing; if the gradient norm is in the hundreds or above → gradient explosion. This is first-hand information for debugging the training process of deep learning models.


Other Extensions