Common Mathematical Symbols

The first hurdle to understanding AI papers and tutorials is often not the mathematics itself, but the pages full of mathematical symbols.

This chapter organizes the most commonly used symbols in AI learning into tables. You can skim through them first, or come back later when you need to look something up:

  • Read in OrderEach symbol comes with an explanation of when it is used; go through them roughly to build an overall impression.

  • Look Up as NeededWhen you get stuck on a symbol while reading papers or tutorials, come back to the table.

Most symbols come with runnable Python examples. Trying them out yourself once is more effective than reading them ten times.


Numbers and Basic Operations

These symbols have been used since elementary school math and are the foundation of all formulas.

SymbolMeaningExample
Addition, subtraction, multiplication, and division
Plus-minus sign, indicating two possible values
Not equal to
Less than, greater than, less than or equal to, greater than or equal to
Approximately equal to
Absolute value, removes the negative sign
Square root
nth root
Square, nth power
Reciprocal (negative exponent)
Factorial, multiply from 1 to n
Infinity

In papers, the multiplication sign is usually omitted or written as a small dot:Represents w multiplied by x; the two notations have exactly the same meaning.


Variables and Letter Conventions

In mathematics, letters are not chosen arbitrarily; different letters have default roles.

LetterDefault roleTypical meaning in AI context
ConstantA fixed, unchanging coefficient
Variablex is the input feature, y is the output label
Weights and biasesParameters that the neural network needs to learn
Parameter setCollective term for all parameters of the model, pronounced theta
QuantityNumber of samples, feature dimension
Pi
Euler's number

When using both x and X, be careful to distinguish:Lowercase usually denotes a single number (scalar), uppercase usually denotes a matrixThis will be elaborated in the lowercase section.

When you seeDon't panic, it's just the conventional symbol for "parameter", andThere is no essential difference; which one the paper author chooses is purely a matter of preference.


Exponents and Logarithms

The most common pair of functions in AI.

Appears in sigmoid and softmax, responsible for "amplifying or compressing" values;Appears in cross-entropy loss, responsible for "turning multiplication into addition".

SymbolMeaningExample
a to the nth power
Any nonzero number to the power of 0
Exponential function with base e
Logarithm of x to base a
Natural logarithm (base e)
In papers, it usually defaults to

The four arithmetic rules must be mastered; they are used repeatedly when deriving loss functions:

The change-of-base formula can convert a logarithm of any base to a natural logarithm; it is commonly used in programming:

Examples

# Calculate log_2(100) using the change-of-base formula
# Python's math.log(x, base) does this internally
import math

x = 100
base = 2

# Direct call
print(math.log(x, base))

# Handwritten change-of-base formula: log_2(100) = ln(100) / ln(2)
print(math.log(x) / math.log(base))

The output of the above code is:

6.643856189774725
6.643856189774725

Why add a small constant ε: The argument of a logarithm must be greater than 0, but early in training the model may output 0 probability; at that timeyou get negative infinity, causing training to collapse.

Therefore, in implementations it is usually written as, whereis a very small positive number, such as 0.0000001.


Function-Related Symbols

A function is the "correspondence rule from input to output"; a neural network is essentially a composition of many layers of functions.

SymbolMeaningExample / Explanation
Function; the result after x is mapped by f
f maps elements of set A to set BDescribe the "type signature" of a function.
Inverse function, undoes the operation of f.
Composite function, first compute g(x) then substitute into f.
The change in x.
The predicted value of y (read as y-hat).To distinguish from the true label y.
sigmoid activation function.

In machine learning,(read as y-hat) specifically refers to the model's predicted value,refers to the true label; the difference between the two is the prediction error.

Composite functions are key to understanding neural networks: a two-layer network is essentially, backpropagation computes gradients layer by layer, relying precisely on the chain rule for differentiating composite functions.

Examples

# Composite function calculation: y = f(g(x))
# Let g(x) = x + 1, f(x) = x^2

def g(x):
    return x + 1          # Inner function

def f(x):
    return x ** 2         # Outer function

# Composition: first compute g(x), then substitute the result into f
print(f(g(3)))            # g(3)=4, f(4)=16
print(f(g(3)) == (3 + 1) ** 2)

Executing the above code outputs:

16
True

Summation and Product

are the most frequently appearing symbols in AI formulas, bar none.

It means: let the subscript i go from the lower bound to the upper bound, and add up each term.

Read as "i from 1 to n, sum over x subscript i". Read a summation expression in three steps:

StepsWhat to look atIn this example
1Which variable is being summed (subscript)i
2Summation range (lower bound to upper bound)1 to n
3What each term looks like

Product symbolThe structure is the same, just replace "addition" with "multiplication":

The two most common AI formulas are built on summation:

Mean squared error loss:

Weighted sum of a neuron:

Examples

# Understand the summation symbol with code: z = sum(w_i * x_i) + b
w = [0.5, -1.0, 2.0]   # Weights, corresponding to w_i in the formula
x = [2.0,  3.0, 1.0]   # Inputs, corresponding to x_i in the formula
b = 0.1                # Bias

# sum_{i=1}^{3} w_i * x_i is exactly the following line
z = sum(wi * xi for wi, xi in zip(w, x)) + b
print(z)

# Expanded verification: 0.5*2 + (-1.0)*3 + 2.0*1 + 0.1
print(0.5*2 + (-1.0)*3 + 2.0*1 + 0.1)

Executing the above code outputs:

0.1
0.1

Sets and Logic

Set symbols describe "which category data belongs to," logic symbols describe "how conditions are combined," and probability and machine learning theory make extensive use of them.

SymbolMeaningExample
belongs to
does not belong to
subset
union
intersection
empty set (contains no elements)
set of real numbers
arbitrary, all (universal quantifier)
exists (existential quantifier)
logical NOT
logical AND, logical OR
implies, implication
equivalent (if and only if)

These symbols most often appear in the theorem and assumption sections of papers, for example,read as "for any epsilon greater than 0, there exists an N".


Superscripts, Subscripts, and Value Ranges

This section resolves two common confusions: what subscripts actually mean, and how to read interval notation.

Subscript: Numbering Variables of the Same Type

Training data has many samples; using x for all of them would be confusing, so subscripts are used to distinguish them:。

When there are two subscripts, such as, it usually means "the i-th row and j-th column", which is exactly an element of a matrix.

A vector is written as(bold), its components are; matrixIts shape is described by "rows × columns", denoted as。

Intervals and Value Ranges

Parentheses mean "endpoints excluded", and square brackets mean "endpoints included".

SymbolReadingMeaning
Open intervalAll numbers between a and b, excluding both endpoints
Closed intervalAll numbers between a and b, including both endpoints
Half-open intervalIncludes b, excludes a
x belongs to the closed intervalx takes values between 0 and 1, endpoints allowed

The range of the sigmoid function is, this open interval means the output is always strictly greater than 0 and less than 1, which is exactly the mathematical reason it can be used as a "probability".


Greek Alphabet Table

The extensive Greek letters in papers are not decoration; each has a default meaning. Knowing them all will make reading papers much smoother.

UppercaseLowercaseNameCommon meanings in AI
alphaLearning rate
betaMomentum coefficient, second-moment decay rate in Adam
gammaDiscount factor (reinforcement learning)
deltaChange, increment
epsilonTiny positive number, exploration probability
thetaModel parameters
lambdaRegularization coefficient
piPolicy (reinforcement learning), pi
sigmaSummation / sigmoid function, standard deviation
phiFeature mapping function
omegaModel weights (some papers)

The same symbol has different meanings in different fields:In the activation function context it is sigmoid, in the statistical context it is standard deviation, and in summation formulas the uppercaseis the summation operator. To understand the meaning, first look at the context—this is the most important habit when reading papers.


Probability and Statistics

The core language of machine learning. The "confidence" output by classification models, the derivation of loss functions, and Bayesian methods are all built on these symbols.

SymbolMeaningExample / Explanation
Probability of event A occurring
Conditional probability: probability of A given BThe vertical bar is read as "given"
Joint probability: A and B occurring simultaneously
Bayes' theoremFoundation of all Bayesian methods
Expectation (mean)
Variance (degree of dispersion)
MeanCenter of the data distribution
Variance (sigma squared)
Follows ... distribution
Normal distribution (Gaussian distribution)Inside the parentheses are mean and variance
Information entropy

Cross-entropy loss is the culmination of these symbols, combining summation, conditional probability, and logarithm:

Breaking down symbol by symbol:is the true label of the i-th sample,is the model's predicted probability,turns the multiplication of probabilities into addition, and the negative sign makes "the smaller the loss, the better".

is the script form of "Normal", referring to the normal distribution, not the letter N. Script symbols、are often used to represent "set-like" objects such as distributions, loss functions, and datasets.


Linear Algebra

The underlying structure of neural networks. Every matrix operation in deep learning frameworks uses these symbols.

SymbolMeaningExample / Explanation
Vector (bold lowercase)
Matrix (bold uppercase)A table of numbers with m rows and n columns
Transpose: swap rows and columnsRow vector becomes column vector
Matrix multiplicationNot element-wise multiplication
Element-wise multiplication (Hadamard product)Multiplying corresponding positions of same-shaped matrices
Norm (the "length" of a vector)
L1 norm: sum of absolute valuesBasis of L1 regularization
Inverse of a matrix
Identity matrixDiagonal entries are 1, all others are 0
Dot product (inner product)

The dot product is the backbone of the attention mechanism: every layer of the transformer computes the dot product between query vectors and key vectors, measuring their "similarity".

Examples

# Symbol reference for dot product and L2 norm
# x · y = sum(x_i * y_i),||x||_2 = sqrt(sum(x_i^2))
import math

x = [1.0, 2.0, 3.0]
y = [4.0, 0.0, -1.0]

# Dot product x · y
dot = sum(xi * yi for xi, yi in zip(x, y))
print(dot)

# L2 norm ||x||_2
norm2 = math.sqrt(sum(xi ** 2 for xi in x))
print(norm2)

# Expand and verify
print(1.0*4.0 + 2.0*0.0 + 3.0*(-1.0))
print(math.sqrt(1 + 4 + 9))

Running the above code produces the following output:

1.0
3.7416573867739413
1.0
3.7416573867739413

Calculus

The essence of training is "adjusting parameters in the opposite direction of the gradient," so the optimization sections in papers cannot do without these symbols.

You don't need to be able to calculate them yet; just be able to recognize each one.

SymbolMeaningExample / Description
Derivative: the rate of change of y with respect to x
Alternative notation for the derivative
Partial derivative: the rate of change of f when only x changesUsed for multivariable functions
Gradient: a vector of all partial derivatives
Integral, the inverse operation of differentiation
Limit: the value as x approaches a

The most important formula in deep learning uses only two rows from the table above: the parameter update rule for gradient descent.

Breaking down each symbol:is the parameter,means "update the left side with the value on the right",is the learning rate,is the gradient of the loss function J with respect to the parameter θ.

In one sentence:new parameter = old parameter − learning rate × gradient, that is, take a small step in the direction that decreases the loss the fastest.

Read as nabla or the "gradient operator", it is not a number but an instruction to take partial derivatives with respect to all variables that follow, much like a higher-order function in programming.


Optimization and Other High-Frequency Symbols

This last section deals with symbols that appear sporadically in papers but are nevertheless critical.

SymbolMeaningExample / Description
the x that maximizes freturns x, not the maximum value itself
the x that minimizes fTraining is finding
the minimum of freturns the value itself
softmax function
function composition
defined asthe left side is defined as the value on the right
proportional to
ReLU activation function
Big O notation: order of growthIgnore constants, only look at scale.

What a classification model does during prediction is exactly: take the class with the highest probability; argmax returns the class index, not the probability value.

andEasily confused:Answer "what is the maximum value",Answer "at which x the maximum is attained." Classification tasks care about the latter.


Study Suggestions

There is no need to memorize symbol tables by rote; methods are more effective than memorization.

First, develop the habit of breaking down symbols one by one: when you get a formula, first ask whether each symbol is an operator, a variable, or a function, then ask what it represents in the current context. After breaking it down, it often ceases to be mysterious.

Second, prioritize thoroughly mastering three things:Summation, exponents and logarithms, subscript notation. They are the prerequisites for reading linear algebra and calculus formulas. This article has dedicated sections for each.

Third, when encountering an unfamiliar symbol, directly search for "symbol name + mathematical meaning", which is faster than slogging through textbooks.

Symbols are the map, not the terrain. True understanding comes from computing concrete examples with them; every Python example in the text is worth running by hand.

other extensions