Common Mathematical Symbols
The first hurdle to understanding AI papers and tutorials is often not the mathematics itself, but the pages full of mathematical symbols.

This chapter organizes the most commonly used symbols in AI learning into tables. You can skim through them first, or come back later when you need to look something up:
-
Read in OrderEach symbol comes with an explanation of when it is used; go through them roughly to build an overall impression.
-
Look Up as NeededWhen you get stuck on a symbol while reading papers or tutorials, come back to the table.
Most symbols come with runnable Python examples. Trying them out yourself once is more effective than reading them ten times.
Numbers and Basic Operations
These symbols have been used since elementary school math and are the foundation of all formulas.
| Symbol | Meaning | Example |
|---|---|---|
| Addition, subtraction, multiplication, and division | ||
| Plus-minus sign, indicating two possible values | ||
| Not equal to | ||
| Less than, greater than, less than or equal to, greater than or equal to | ||
| Approximately equal to | ||
| Absolute value, removes the negative sign | ||
| Square root | ||
| nth root | ||
| Square, nth power | ||
| Reciprocal (negative exponent) | ||
| Factorial, multiply from 1 to n | ||
| Infinity |
In papers, the multiplication sign is usually omitted or written as a small dot:Represents w multiplied by x; the two notations have exactly the same meaning.
Variables and Letter Conventions
In mathematics, letters are not chosen arbitrarily; different letters have default roles.
| Letter | Default role | Typical meaning in AI context |
|---|---|---|
| Constant | A fixed, unchanging coefficient | |
| Variable | x is the input feature, y is the output label | |
| Weights and biases | Parameters that the neural network needs to learn | |
| Parameter set | Collective term for all parameters of the model, pronounced theta | |
| Quantity | Number of samples, feature dimension | |
| Pi | ||
| Euler's number |
When using both x and X, be careful to distinguish:Lowercase usually denotes a single number (scalar), uppercase usually denotes a matrixThis will be elaborated in the lowercase section.
When you seeDon't panic, it's just the conventional symbol for "parameter", andThere is no essential difference; which one the paper author chooses is purely a matter of preference.
Exponents and Logarithms
The most common pair of functions in AI.
Appears in sigmoid and softmax, responsible for "amplifying or compressing" values;Appears in cross-entropy loss, responsible for "turning multiplication into addition".
| Symbol | Meaning | Example |
|---|---|---|
| a to the nth power | ||
| Any nonzero number to the power of 0 | ||
| Exponential function with base e | ||
| Logarithm of x to base a | ||
| Natural logarithm (base e) | ||
| In papers, it usually defaults to |
The four arithmetic rules must be mastered; they are used repeatedly when deriving loss functions:
The change-of-base formula can convert a logarithm of any base to a natural logarithm; it is commonly used in programming:
Examples
# Python's math.log(x, base) does this internally
import math
x = 100
base = 2
# Direct call
print(math.log(x, base))
# Handwritten change-of-base formula: log_2(100) = ln(100) / ln(2)
print(math.log(x) / math.log(base))
The output of the above code is:
6.643856189774725 6.643856189774725
Why add a small constant ε: The argument of a logarithm must be greater than 0, but early in training the model may output 0 probability; at that timeyou get negative infinity, causing training to collapse.
Therefore, in implementations it is usually written as, whereis a very small positive number, such as 0.0000001.
Function-Related Symbols
A function is the "correspondence rule from input to output"; a neural network is essentially a composition of many layers of functions.
| Symbol | Meaning | Example / Explanation |
|---|---|---|
| Function; the result after x is mapped by f | ||
| f maps elements of set A to set B | Describe the "type signature" of a function. | |
| Inverse function, undoes the operation of f. | ||
| Composite function, first compute g(x) then substitute into f. | ||
| The change in x. | ||
| The predicted value of y (read as y-hat). | To distinguish from the true label y. | |
| sigmoid activation function. |
In machine learning,(read as y-hat) specifically refers to the model's predicted value,refers to the true label; the difference between the two is the prediction error.
Composite functions are key to understanding neural networks: a two-layer network is essentially, backpropagation computes gradients layer by layer, relying precisely on the chain rule for differentiating composite functions.
Examples
# Let g(x) = x + 1, f(x) = x^2
def g(x):
return x + 1 # Inner function
def f(x):
return x ** 2 # Outer function
# Composition: first compute g(x), then substitute the result into f
print(f(g(3))) # g(3)=4, f(4)=16
print(f(g(3)) == (3 + 1) ** 2)
Executing the above code outputs:
16 True
Summation and Product
are the most frequently appearing symbols in AI formulas, bar none.
It means: let the subscript i go from the lower bound to the upper bound, and add up each term.
Read as "i from 1 to n, sum over x subscript i". Read a summation expression in three steps:
| Steps | What to look at | In this example |
|---|---|---|
| 1 | Which variable is being summed (subscript) | i |
| 2 | Summation range (lower bound to upper bound) | 1 to n |
| 3 | What each term looks like |
Product symbolThe structure is the same, just replace "addition" with "multiplication":
The two most common AI formulas are built on summation:
Mean squared error loss:
Weighted sum of a neuron:
Examples
w = [0.5, -1.0, 2.0] # Weights, corresponding to w_i in the formula
x = [2.0, 3.0, 1.0] # Inputs, corresponding to x_i in the formula
b = 0.1 # Bias
# sum_{i=1}^{3} w_i * x_i is exactly the following line
z = sum(wi * xi for wi, xi in zip(w, x)) + b
print(z)
# Expanded verification: 0.5*2 + (-1.0)*3 + 2.0*1 + 0.1
print(0.5*2 + (-1.0)*3 + 2.0*1 + 0.1)
Executing the above code outputs:
0.1 0.1
Sets and Logic
Set symbols describe "which category data belongs to," logic symbols describe "how conditions are combined," and probability and machine learning theory make extensive use of them.
| Symbol | Meaning | Example |
|---|---|---|
| belongs to | ||
| does not belong to | ||
| subset | ||
| union | ||
| intersection | ||
| empty set (contains no elements) | ||
| set of real numbers | ||
| arbitrary, all (universal quantifier) | ||
| exists (existential quantifier) | ||
| logical NOT | ||
| logical AND, logical OR | ||
| implies, implication | ||
| equivalent (if and only if) |
These symbols most often appear in the theorem and assumption sections of papers, for example,read as "for any epsilon greater than 0, there exists an N".
Superscripts, Subscripts, and Value Ranges
This section resolves two common confusions: what subscripts actually mean, and how to read interval notation.
Subscript: Numbering Variables of the Same Type
Training data has many samples; using x for all of them would be confusing, so subscripts are used to distinguish them:。
When there are two subscripts, such as, it usually means "the i-th row and j-th column", which is exactly an element of a matrix.
A vector is written as(bold), its components are; matrixIts shape is described by "rows × columns", denoted as。
Intervals and Value Ranges
Parentheses mean "endpoints excluded", and square brackets mean "endpoints included".
| Symbol | Reading | Meaning |
|---|---|---|
| Open interval | All numbers between a and b, excluding both endpoints | |
| Closed interval | All numbers between a and b, including both endpoints | |
| Half-open interval | Includes b, excludes a | |
| x belongs to the closed interval | x takes values between 0 and 1, endpoints allowed |
The range of the sigmoid function is, this open interval means the output is always strictly greater than 0 and less than 1, which is exactly the mathematical reason it can be used as a "probability".
Greek Alphabet Table
The extensive Greek letters in papers are not decoration; each has a default meaning. Knowing them all will make reading papers much smoother.
| Uppercase | Lowercase | Name | Common meanings in AI |
|---|---|---|---|
| alpha | Learning rate | ||
| beta | Momentum coefficient, second-moment decay rate in Adam | ||
| gamma | Discount factor (reinforcement learning) | ||
| delta | Change, increment | ||
| epsilon | Tiny positive number, exploration probability | ||
| theta | Model parameters | ||
| lambda | Regularization coefficient | ||
| pi | Policy (reinforcement learning), pi | ||
| sigma | Summation / sigmoid function, standard deviation | ||
| phi | Feature mapping function | ||
| omega | Model weights (some papers) |
The same symbol has different meanings in different fields:In the activation function context it is sigmoid, in the statistical context it is standard deviation, and in summation formulas the uppercaseis the summation operator. To understand the meaning, first look at the context—this is the most important habit when reading papers.
Probability and Statistics
The core language of machine learning. The "confidence" output by classification models, the derivation of loss functions, and Bayesian methods are all built on these symbols.
| Symbol | Meaning | Example / Explanation |
|---|---|---|
| Probability of event A occurring | ||
| Conditional probability: probability of A given B | The vertical bar is read as "given" | |
| Joint probability: A and B occurring simultaneously | ||
| Bayes' theorem | Foundation of all Bayesian methods | |
| Expectation (mean) | ||
| Variance (degree of dispersion) | ||
| Mean | Center of the data distribution | |
| Variance (sigma squared) | ||
| Follows ... distribution | ||
| Normal distribution (Gaussian distribution) | Inside the parentheses are mean and variance | |
| Information entropy |
Cross-entropy loss is the culmination of these symbols, combining summation, conditional probability, and logarithm:
Breaking down symbol by symbol:is the true label of the i-th sample,is the model's predicted probability,turns the multiplication of probabilities into addition, and the negative sign makes "the smaller the loss, the better".
is the script form of "Normal", referring to the normal distribution, not the letter N. Script symbols、are often used to represent "set-like" objects such as distributions, loss functions, and datasets.
Linear Algebra
The underlying structure of neural networks. Every matrix operation in deep learning frameworks uses these symbols.
| Symbol | Meaning | Example / Explanation |
|---|---|---|
| Vector (bold lowercase) | ||
| Matrix (bold uppercase) | A table of numbers with m rows and n columns | |
| Transpose: swap rows and columns | Row vector becomes column vector | |
| Matrix multiplication | Not element-wise multiplication | |
| Element-wise multiplication (Hadamard product) | Multiplying corresponding positions of same-shaped matrices | |
| Norm (the "length" of a vector) | ||
| L1 norm: sum of absolute values | Basis of L1 regularization | |
| Inverse of a matrix | ||
| Identity matrix | Diagonal entries are 1, all others are 0 | |
| Dot product (inner product) |
The dot product is the backbone of the attention mechanism: every layer of the transformer computes the dot product between query vectors and key vectors, measuring their "similarity".
Examples
# x · y = sum(x_i * y_i),||x||_2 = sqrt(sum(x_i^2))
import math
x = [1.0, 2.0, 3.0]
y = [4.0, 0.0, -1.0]
# Dot product x · y
dot = sum(xi * yi for xi, yi in zip(x, y))
print(dot)
# L2 norm ||x||_2
norm2 = math.sqrt(sum(xi ** 2 for xi in x))
print(norm2)
# Expand and verify
print(1.0*4.0 + 2.0*0.0 + 3.0*(-1.0))
print(math.sqrt(1 + 4 + 9))
Running the above code produces the following output:
1.0 3.7416573867739413 1.0 3.7416573867739413
Calculus
The essence of training is "adjusting parameters in the opposite direction of the gradient," so the optimization sections in papers cannot do without these symbols.
You don't need to be able to calculate them yet; just be able to recognize each one.
| Symbol | Meaning | Example / Description |
|---|---|---|
| Derivative: the rate of change of y with respect to x | ||
| Alternative notation for the derivative | ||
| Partial derivative: the rate of change of f when only x changes | Used for multivariable functions | |
| Gradient: a vector of all partial derivatives | ||
| Integral, the inverse operation of differentiation | ||
| Limit: the value as x approaches a |
The most important formula in deep learning uses only two rows from the table above: the parameter update rule for gradient descent.
Breaking down each symbol:is the parameter,means "update the left side with the value on the right",is the learning rate,is the gradient of the loss function J with respect to the parameter θ.
In one sentence:new parameter = old parameter − learning rate × gradient, that is, take a small step in the direction that decreases the loss the fastest.
Read as nabla or the "gradient operator", it is not a number but an instruction to take partial derivatives with respect to all variables that follow, much like a higher-order function in programming.
Optimization and Other High-Frequency Symbols
This last section deals with symbols that appear sporadically in papers but are nevertheless critical.
| Symbol | Meaning | Example / Description |
|---|---|---|
| the x that maximizes f | returns x, not the maximum value itself | |
| the x that minimizes f | Training is finding | |
| the minimum of f | returns the value itself | |
| softmax function | ||
| function composition | ||
| defined as | the left side is defined as the value on the right | |
| proportional to | ||
| ReLU activation function | ||
| Big O notation: order of growth | Ignore constants, only look at scale. |
What a classification model does during prediction is exactly: take the class with the highest probability; argmax returns the class index, not the probability value.
andEasily confused:Answer "what is the maximum value",Answer "at which x the maximum is attained." Classification tasks care about the latter.
Study Suggestions
There is no need to memorize symbol tables by rote; methods are more effective than memorization.
First, develop the habit of breaking down symbols one by one: when you get a formula, first ask whether each symbol is an operator, a variable, or a function, then ask what it represents in the current context. After breaking it down, it often ceases to be mysterious.
Second, prioritize thoroughly mastering three things:Summation, exponents and logarithms, subscript notation. They are the prerequisites for reading linear algebra and calculus formulas. This article has dedicated sections for each.
Third, when encountering an unfamiliar symbol, directly search for "symbol name + mathematical meaning", which is faster than slogging through textbooks.
other extensionsSymbols are the map, not the terrain. True understanding comes from computing concrete examples with them; every Python example in the text is worth running by hand.