Derivative Introduction -- Instantaneous Rate of Change and Tangent Line

Derivative = the instantaneous rate of change of a function at a point = the slope of the tangent line at that point.

Understanding the derivative helps you understand why gradient descent moves in the opposite direction.


From Average Rate of Change to Instantaneous Rate of Change

Average rate of change (secant slope) → Δx gets smaller and smaller → secant becomes tangent →Derivative = slope of the tangent line。

The Sign of the Derivative Reveals Function Behavior

f'(x) > 0

The function is rising here

The tangent line slopes upward to the right

f'(x) < 0

The function is falling here

The tangent line slopes downward to the right

f'(x) = 0

Horizontal — possibly an extremum

The tangent line is horizontal

The intuition behind gradient descent lies in the sign of the derivative:

f'(x) > 0 → the function is rising → move in the opposite direction. f'(x) < 0 → the function is falling → keep moving in this direction.

This is the root of the 'negative gradient direction'.


Real-life Example

Speedometer = derivative of the position function

The whole trip was 120 km and took 2 hours → average speed 60 km/h.

But the instantaneous speed on the dashboard keeps changing. This 'speed at this moment' is the derivative of the distance function with respect to time.


Mathematical Definition

\[ f'(x) = \lim_{\Delta x \to 0} \frac{f(x + \Delta x) - f(x)}{\Delta x} \]

If this limit exists, f is differentiable at x.


Python Hands-on Practice

Example

import numpy as np

def numerical_derivative(f, x, h=1e-5):
    return (f(x + h) - f(x - h)) / (2 * h)

f = lambda x: x**2

# f'(2) = 2*2 = 4
print(f"f'(2) value={numerical_derivative(f, 2):.6f}, theoretical=4")

# Derivative at each point
for x in [-2, -1, 0, 1, 2]:
    d = numerical_derivative(f, x)
    dir = falling if d < 0 else (rising if d > 0 else horizontal)
    print(f"x={x:2d}, f'(x)={d:5.1f}, {dir}")
f'(2) 数值=4.000000, 理论=4
x=-2, f'(x)= -4.0, 下降
x=-1, f'(x)= -2.0, 下降
x= 0, f'(x)=  0.0, 水平
x= 1, f'(x)=  2.0, 上升
x= 2, f'(x)=  4.0, 上升

Application Scenarios in AI

One-dimensional Prototype of Gradient Descent

For a univariate function, gradient descent reduces to \( x_{new} = x_{old} - \eta \cdot f'(x_{old}) \). The derivative tells you which direction the function value decreases — if the derivative is positive, move left; if negative, move right.

The derivative of the activation function determines the efficiency of backpropagation

The derivative of ReLU is 1 when x > 0 and 0 when x ≤ 0. It is extremely fast to compute, and the gradient does not decay in the positive region—this is the key reason ReLU is more suitable for deep networks than Sigmoid. The maximum derivative of Sigmoid is only 0.25, and after being multiplied across multiple layers, the gradient decays exponentially (vanishing gradient).

The derivative of the loss function is zero at the optimal solution

When the model converges to a local optimum, the partial derivatives of the loss function with respect to all parameters are close to 0—gradient descent naturally stops. This is a signal of training convergence: the gradient norm ||\nabla J|| \approx 0.


Other extensions