Functions and Limits -- The Starting Line of Calculus
Calculus is the mathematical engine of the AI training process. Before learning derivatives, first review the concept of functions, then understand "limits" — the core idea of calculus.
What is a limit?
A limit describes "what value f(x) approaches as x gets infinitely close to but not equal to a certain point."
For example: \( f(x) = \frac{x^2-1}{x-1} \) is undefined at x=1. But when x approaches 1, f(x) approaches 2.
Left-hand limit
x → 1⁻ (approach from the left)
f(0.999) → 2
Limit
Left and right limits are equal → the limit exists
\( \lim_{x\to 1} f(x) = 2 \)
Right-hand limit
x → 1⁺ (approach from the right)
f(1.001) → 2
Continuous = no breaks
f(x) is continuous at a: \( \lim_{x\to a} f(x) = f(a) \)— The limit at that point equals the function value at that point.
Most loss functions (MSE, cross-entropy) are continuous and differentiable — this is the prerequisite for optimization using gradient descent.
Real-life examples
Instantaneous velocity
You want to know the velocity at the exact moment of the 3rd second. But the time interval of an "instant" is 0, and the displacement is also 0, so 0/0 is meaningless.
The calculus solution: take increasingly shorter time intervals and see what value the average velocity approaches. The approached value is the instantaneous velocity — that is, the derivative.
Mathematical definition
\[ \lim_{x \to a} f(x) = L \]Limit operation rules: the limit of a sum, difference, product, or quotient = the sum, difference, product, or quotient of the limits (provided the denominator limit is not 0).
Python hands-on practice
Example
def f(x):
return (x**2 - 1) / (x - 1)
print("=== lim_{x->1} (x^2-1)/(x-1) = 2 ===")
# Approach from both the left and right
for delta in [0.1, 0.01, 0.001, 0.0001]:
print(f"x = {1-delta:.4f}, f(x) = {f(1-delta):.6f}")
print(f"x = {1+delta:.4f}, f(x) = {f(1+delta):.6f}")
print("→ No matter which side you approach from, f(x) approaches 2")
=== lim_{x->1} (x^2-1)/(x-1) = 2 ===
x = 0.9000, f(x) = 1.900000
x = 1.1000, f(x) = 2.100000
x = 0.9900, f(x) = 1.990000
x = 1.0100, f(x) = 2.010000
→ 无论从哪边逼近,f(x)都趋近于 2
Application scenarios in AI
The definition of a gradient depends on limits
The definition of the derivative itself is a limit: \( f'(x) = \lim_{h \to 0} \frac{f(x+h)-f(x)}{h} \). Although gradient computation in backpropagation uses analytic differentiation rather than numerical limits, the idea of limits is key to understanding that "the learning rate cannot be too large"—a learning rate that is too large is equivalent to approximating the derivative with a value of \(h\) that is too large.
Continuity requirements of loss functions
Gradient descent requires the loss function to be differentiable almost everywhere. MSE and cross-entropy are both smooth and continuous—this is the prerequisite for their becoming standard loss functions. Hinge Loss (used in SVM) is non-differentiable at the turning point and needs to be replaced with a subgradient.
Choice of activation functions
ReLU is not differentiable at x=0 (the left and right derivatives are unequal), but in practice the subgradient 0 is used instead, and the model can still train normally. LeakyReLU and GELU are differentiable everywhere, and their gradient flow is smoother.
Other extensions