Finding the Minimum of a Function Using Derivatives
Use numerical differences to estimate the derivative, and move step by step in the opposite direction of the derivative, to get an intuitive feel that "the derivative guides the search direction."
After finishing this case study, you will understand:The derivative tells you which direction makes the function value increase; going in the opposite direction makes it decrease—this is the prototype of gradient descent.
A Real-Life Introduction
Going Down the Mountain in Dense Fog
In a dense fog where you can't see your hand in front of your face, you need to walk from the mountaintop to the valley. You cannot see the overall terrain; you can only feel the slope under your feet. Your strategy is simple: feel the direction of the steepest slope → take a small step in the steepest downhill direction → stop → feel again → take another small step.
"Feeling the slope" is computing the derivative, and "taking a small step" is updating the position using the opposite direction of the derivative.This method does not require knowing the shape of the entire mountain; it only requires knowing the slope under your feet.
Intuitive Understanding
For the function \(f(x) = (x-3)^2 + 2\), the minimum is at \(x=3\), and the minimum value is \(f(3) = 2\).
Starting from \(x=10\), how do you get to \(x=3\)? You don't need to solve an equation; you just need a strategy:
- If the derivative \(f'(x) > 0\) (the function is increasing), go left (decrease x).
- If the derivative \(f'(x) < 0\) (the function is decreasing), go right (increase x).
- The larger the absolute value of the derivative, the steeper the slope, so you can take a larger step.
Mathematical Definition
Numerical Derivative (Central Difference)
\[ f'(x) \approx \frac{f(x+h) - f(x-h)}{2h} \]There is no need to derive the differentiation formula. Use the difference in function values at two very close points divided by the distance to get an approximation of the derivative. Take \(h\) as an extremely small number, such as \(10^{-5}\).
Naive Downhill Algorithm
\[ x_{t+1} = x_t - \eta \cdot f'(x_t) \]Here \(\eta\) is the step size (learning rate), which controls how far each step goes.
Python Hands-On Practice
Example
# Objective function: f(x) = (x-3)^2 + 2, minimum at x=3
def f(x):
return (x - 3) ** 2 + 2
# Numerical derivative — using central difference, not relying on symbolic differentiation
def numerical_derivative(f, x, h=1e-5):
"""Estimate the derivative of f at x."""
return (f(x + h) - f(x - h)) / (2 * h)
# Naive Descent Algorithm: Move in the Opposite Direction of the Derivative
x = 10.0 # Starting point, deliberately chosen far from the optimal point
step_size = 0.1 # Step size for each move
history = [x]
for i in range(50):
grad = numerical_derivative(f, x)
x = x - step_size * grad # If the derivative is positive, move left; if negative, move right
history.append(x)
print("EXAMPLE Descent Algorithm Result:")
print(f"Starting point x = {history[0]:.1f}")
print(f"End point x = {history[-1]:.4f} (true minimum point x=3)")
print(f"End point f(x) = {f(history[-1]):.4f} (true minimum f(3)=2)")
print("\nEXAMPLE first 10-step movement trajectory:")
print(f"{'Step':<6} {'x':<10} {'f(x)':<10} {'Derivative':<10} {'Direction'}")
for i, xi in enumerate(history[:10]):
grad = numerical_derivative(f, xi)
direction = Left if grad > 0 else Right
print(f"{i:<6} {xi:<10.4f} {f(xi):<10.4f} {grad:<10.4f} {direction}")
EXAMPLE 下山算法结果: 起点 x = 10.0 终点 x = 3.0000 (真实最小值点 x=3) 终点 f(x) = 2.0000 (真实最小值 f(3)=2) EXAMPLE 前 10 步移动轨迹: 步数 x f(x) 导数 方向 0 10.0000 51.0000 14.0000 往左 1 8.6000 33.3600 11.2000 往左 2 7.4800 22.0704 8.9600 往左 3 6.5840 14.8042 7.1680 往左 4 5.8672 10.1942 5.7344 往左
Application Scenarios in AI
| Scenario | Connection to this case |
|---|---|
| Gradient descent training | The essence of neural network training: using the partial derivative of the loss function with respect to each parameter to guide the parameter update direction |
| Numerical gradient checking | When debugging backpropagation, use numerical gradients to verify the correctness of analytical gradients—this is exactly the central difference method in this case. |
| Hyperparameter search | Tuning methods such as Bayesian optimization also search for the minimum on an unknown function surface along the direction of the "derivative". |