Gradient -- points in the direction of the fastest increase of the function.
The gradient points in the direction in which the function value increases fastest, so gradient descent goes in the negative gradient direction.
Gradient is the core concept of the entire AI training process — every update step taken by the optimizer relies on the gradient.
Concept Analysis
Gradient = a vector composed of partial derivatives
\[ \nabla f(x, y) = \left( \frac{\partial f}{\partial x}, \frac{\partial f}{\partial y} \right) \]Gradient Direction
Points in the direction of steepest ascent
Want to go uphill → follow the gradient
Negative gradient direction
Points in the direction of steepest descent
Want to go downhill → follow the negative gradient
Gradient Magnitude
The steepness of that direction
Gradient is zero → extremum or saddle point
The negative sign in the update formula \( \theta = \theta - \eta \nabla J \) means "move in the opposite direction (downhill)."
If you forget the negative sign, it becomes gradient ascent, and the loss will get larger and larger.
Everyday Examples
Descending the mountain on a foggy day
You are standing on a mountain shrouded in fog. To go down as quickly as possible, you feel which direction under your feet is steepest, then take a step in the steepest downhill direction.
Repeat continuously: feel → step → feel → step. This is the physical intuition of gradient descent.
The gradient tells you the "steepest direction," and the negative sign tells you to "go down."
Mathematical Definition
\[ \nabla f(\mathbf{x}) = \begin{bmatrix} \frac{\partial f}{\partial x_1} \\ \vdots \\ \frac{\partial f}{\partial x_n} \end{bmatrix} \]direction导number沿 u direction:\( \nabla f \cdot \mathbf{u} = \|\nabla f\| \cos\theta \)。When u and梯degreesametowardwhenmaximum。
Hands-on Python practice
Example
def f(x, y):
return x**2 + y**2 # Bowl-shaped function
def grad(x, y):
return np.array([2*x, 2*y])
pt = np.array([1.5, 1.0])
g = grad(*pt)
print(fAt (1.5, 1.0): gradient = {g})
print(fGradient norm = {np.linalg.norm(g):.2f})
print(fThe gradient points away from the origin (ascent), and the negative gradient points toward the origin (descent))
Interactive 3D visualization of the loss surface
Below we use Plotly.js to draw the 3D surface of \( f(x,y)=x^2+y^2 \) to intuitively feel the gradient direction:
Application scenarios in AI
Core input for gradient descent training
Every parameter update \( \theta = \theta - \eta \nabla_\theta J \) depends on the gradient. The gradient tells each parameter "increase or decrease, and by how much." In PyTorch, loss.backward() computes the gradient, and optimizer.step() uses the gradient to update the parameters.
Vanishing gradients and exploding gradients
In deep networks, during backpropagation the gradient may become increasingly small (vanishing, causing shallow layers to learn nothing) or increasingly large (exploding, causing parameter updates to be too large and diverge). ResNet's residual connections (identity shortcut) are designed to alleviate gradient vanishing—giving the gradient a "highway" that goes straight to the shallow layers.
Gradient Clipping
When training RNNs and large language models, gradient clipping is commonly applied: if the L2 norm of the gradient exceeds a threshold, it is scaled down to the threshold size. This is a practical technique to prevent gradient explosion; the standard practice in GPT training is clip_grad_norm_=1.0.
Practical significance of visualizing gradients
Tools like TensorBoard and W&B can visualize the distribution of gradients in each layer. If the gradient of a layer is almost zero → gradient vanishing; if the gradient norm is in the hundreds or above → gradient explosion. This is first-hand information for debugging the training process of deep learning models.
Other Extensions