Functions and Graphs

In this chapter, we will visually display the graphs of common functions; understanding the shape of a graph is more important than memorizing formulas.

Loss function surfaces and activation functions are essentially variants of these basic shapes. Each function comes with a switchable interactive graph; switching and dragging by hand is far more effective than looking at static illustrations.


Why specifically learn to "read graphs"

Mathematical formulas describe "rules," while graphs describe "what the rules look like."

Many core problems in deep learning are essentially "curve or surface shape" problems:

Deep learning problemUnderlying shape problem
Why some loss functions are easier to optimizeSee whether its surface is "bowl-shaped" (convex function)
Why sigmoid causes slower trainingSee how "flat" its graph is at both ends
Why ReLU became the default choiceSee how "simple" its graph is

Looking at each function with questions like "What shape is this? How are the ends? How is the middle?" will help you remember it better than rote memorizing formulas.


Five basic function graphs

Linear function: y = kx + b

Its shape is a straight line.k is the slope, determining the degree and direction of tilt (k>0 slopes up to the right, k<0 slopes down to the right);b is the intercept, determining the point where the line intersects the y-axis.

Both the domain and range are all real numbers, and the rate of change (slope) is the same everywhere; this is the meaning of the word "linear."

Role in AI: the most basic operation of a neuronIt is precisely a linear function in high-dimensional space. Weights control the direction and degree of tilt, and the bias controls the overall translation.

Quadratic function: y = ax² + bx + c

Its shape is a parabola.Opening upward (bowl-shaped), has a minimum value;Opening downward (inverted bowl), has a maximum value.

The vertex is the turning point of the curve and can be found by completing the square:The graph is symmetric about the vertical line through the vertex.

Role in AI: the simplest convex function looks like this. Mean squared error (MSE) approximates this "bowl-shaped" surface in many simplified scenarios—gradient descent easily finds the lowest point on a bowl-shaped surface because no matter which direction you descend, you head toward the same valley floor.

Exponential function: y = aˣ (a>0 and a≠1)

The shape is a curve that rises with accelerating speed (when a>1) or continuously decays (when 0<a<1); it never touches the x-axis, which is its horizontal asymptote.

The domain is all real numbers, and the range isWhen the input grows linearly, the output grows "explosively"—this is the core property of the exponential function.

Role in AI: in softmax,and the exponential decay in the momentum term of the Adam optimizer both exploit the properties of "amplifying differences" or "gradually forgetting".

Logarithmic function: y = logₐx (a>0 and a≠1)

It is exactly the "mirror image" of the exponential function (with the line y=x as the axis of symmetry). The curve rises rapidly from x=0, then growth becomes slower and slower.

The domain is, and the range is all real numbers. The larger the input, the slower the growth—completely the opposite of the exponential function.

Role in AI: in cross-entropy loss,it precisely leverages the logarithm's shape—"extremely steep near 0, flattening near 1"—the more confidently wrong the model is, the heavier the penalty.

Trigonometric functions: y = sin(x), y = cos(x)

The shape is a wave that repeats continuously (periodicity) and always oscillates between -1 and 1 (boundedness). sin(x) starts from the origin and goes upward; cos(x) reaches its maximum value 1 at x=0.

The period is, and the range is。

Role in AI: Transformer positional encoding directly overlays sin/cos waveforms of different frequencies to generate a unique yet regular "coordinate" for each position in the sequence (expanded in Chapter 3).

Interactive demo: gallery of five basic function graphs
Click the button to switch functions; the info card below updates shape characteristics and AI roles simultaneously.

Three key AI activation functions

Activation functions are the "source of nonlinearity" for each layer of a neural network, and their shapes directly determine how easy or difficult training is.

sigmoid:σ(x) = 1 / (1 + e⁻ˣ)

The shape is an S-shaped curve that slowly climbs from 0 to 1, rising fastest in the middle (near x=0), and flattening out at both ends.

The value range is an open interval, often interpreted as 'probability'.

The biggest side effect: the slope at both ends approaches 0.When the absolute value of the input is large (e.g., x=10 or x=-10), the sigmoid curve is nearly flat, and the gradient is nearly zero.

In deep networks, if many layers all use sigmoid, the gradient gets weakened layer by layer by these "flat regions" during backpropagation, and ultimately cannot reach the earlier layers—this is...The vanishing gradient problemThe most intuitive geometric origin.

ReLU:f(x) = max(0, x)

The shape is a piecewise linear line: when x<0 it hugs the x-axis (output is always 0), and when x≥0 it is a straight line with slope 1.

The derivative is always 0 on the negative half-axis and always 1 on the positive half-axis. There's no "gradually flattening" process like sigmoid, and the computation is extremely simple (only one comparison needed).

The gradient on the positive half-axis is always 1 and does not decay as the input grows larger, greatly alleviating the vanishing gradient problem. This is the intuitive reason why ReLU became the default activation function.

italsoYesownProblem:negative半轴梯degreeHengis 0,possible导致some神经元"死掉"No再Update--这fromimage形stateTopalsoabilitystraightJie看出come。 -> It also has its own problems: the gradient on the negative half-axis is always 0, which may cause some neurons to "die" and stop updating—this can be directly seen from the shape of the image.

tanh:f(x) = (eˣ - e⁻ˣ) / (eˣ + e⁻ˣ)

An S-shaped curve very similar to sigmoid, but centered at the origin, with a range of...。

Compared to sigmoid, tanh's output is centered at 0. If the next layer's input mean is close to 0, training tends to be more stable—this is why tanh is more popular in certain scenarios.

But it doesn't solve the fundamental problem of "flattening at both ends," which is also the background for why ReLU-family functions later became popular.

Interactive Demo: Activation Function and Its Derivative
Switch the activation function. The solid line is the function itself, and the dashed line is its derivative (slope). Note the peak of the derivative curve and the decay at both ends.
sigmoid of导numbermaximum valueOnly 0.25, andIn |x| > 4 when几乎is 0--multi-layerconnectmultiplyafter梯degreerapidlyeliminateloss,This is"梯degreeeliminate失"of几whatSource。

Why cross-entropy uses -log(p)

In binary classification problems, the probability that the model predicts the "correct class" is..., the cross-entropy loss is。

Intuitively, you might think: useWouldn't the loss be simpler? When the two curves are placed together, the difference is immediately visible.

Interactive demo: penalty comparison of -log(p) vs 1-p
Drag滑blockchangeModelPredicted probability p,observationtwo种损失Give出of惩罚Poordifferent。When p 很small(Model自信ground犯错)when,-log(p) of惩罚fargreater than 1-p。
Blue is -log(p), green is 1-p. The former shoots to infinity as p → 0, while the latter reaches at most 1.

Examples

# Work it out by hand: vanishing gradients and the -log(p) penalty
import math

# ---- Part 1: The sigmoid derivative decays as |x| increases ----
def sigmoid(x):
    return 1 / (1 + math.exp(-x))

def sigmoid_grad(x):
    s = sigmoid(x)
    return s * (1 - s)     # Derivative formula of sigmoid

print(sigmoid_grad(0))    # At the peak: 0.25, already the maximum value of the sigmoid derivative
print(sigmoid_grad(5))    # When |x|=5: less than 0.007 remains
print(sigmoid_grad(10))   # When |x|=10: about 0.0000456, the gradient almost vanishes

# ---- Part 2: The growth of the -log(p) penalty ----
for p in [1.0, 0.5, 0.1, 0.01]:
    print(p, -math.log(p))   # The smaller p, the steeper the penalty; at p=0.01, the penalty has already reached 4.6

Running the above code produces the following output:

0.25
0.006648056670790155
4.5395807735957664e-05
1.0 0.0
0.5 0.6931471805599453
0.1 2.302585092994046
0.01 4.605170185988091

Two numbers are worth remembering: the peak of the sigmoid derivative is only0.25; when p drops from 0.5 to 0.01,the penalty rises from 0.69 to 4.6.


Three scenarios for applying "shape intuition"

ScenarioWhat shape to look forCriterion
Judging whether the loss function is easy to optimizeWhether the surface is close to 'bowl-shaped' (convex)The closer to a bowl shape, the easier it is for gradient descent to find the global minimum; a bumpy (non-convex) surface may get stuck in a local optimum
Judging whether the activation function slows down trainingWhether the two ends of the graph 'flatten out'The flatter, the more obvious the vanishing gradient problem
Judging whether a function is suitable for probability outputWhether the range is boundedOnly when the range is compressed into an interval like (0,1) can it be interpreted as a probability; this is the direct reason why sigmoid and softmax are chosen

Exercise: Without looking at formulas, judge by intuition alone—between ReLU and sigmoid, which one is more suitable for the final layer of the network for 'binary classification probability output'? Why?

Click to view the answer

sigmoid is more suitable because its range is, which can naturally be interpreted as 'the probability of belonging to a certain class'.

ReLU's range is, has no upper bound, so it cannot be directly used as a probability.


Chapter summary

FunctionShape keywordsCorresponding AI concept
Linear functionStraight line, same slope everywhereNeuron's weighted sum wx+b
Quadratic functionBowl-shaped / inverted bowl-shapedLoss function surface (convex function intuition)
Exponential functionAccelerated rise or decaysoftmax, momentum term
Logarithmic functionGrowth is getting slower and slower.Cross-entropy loss
Trigonometric functionsPeriodic wavesPositional encoding
sigmoidS-shaped, flattened at both ends.Probability output, the source of vanishing gradients.
ReLUPolyline, negative half-axis returns to zero.Default activation function
tanhS-shaped, centered at 0.Scenarios requiring "zero-mean output"
Other extensions