Geometry and Trigonometry
This chapter builds intuition for vectors as 'direction + length', develops an understanding of basic trigonometric functions, and uses the spatial intuition of coordinate systems as a stepping stone for learning 'vector spaces' in linear algebra.
Word embeddings, similarity calculations in recommendation systems, and positional encodings in Transformers are all built on the geometric intuition from this chapter.
Plane Vectors: Direction + Length
What is a Vector?
An ordinary number (scalar) carries only one piece of information: 'magnitude', for example, 'the temperature is 25 degrees'.
Vectorcarries two pieces of information simultaneously:DirectionandLength (Magnitude)。
In a plane rectangular coordinate system, a vector is written in coordinate form, which represents the position reached by 'starting from the origin, moving x units along the x direction and moving y units along the y direction'. The arrow itself encodes direction and length.
Vector Length (Magnitude)
This is the Pythagorean theorem: the length of a vector is the hypotenuse of a right triangle formed by its components along the x-axis and y-axis.
Vector Addition and Scalar Multiplication
| Operation | Algebraic Rule | Geometric Meaning |
|---|---|---|
| Addition | Tip-to-tail: connect the start of the second vector to the end of the first vector | |
| Scalar Multiplication | Only changes the length (reverses direction when k is negative), without changing the line along which the direction lies |
Vector Dot Product: Measuring How Close Two Directions Are
whereis the angle between the two vectors. The sign of the dot product has a clear geometric meaning:
| Sign of Dot Product | Angle Range | Geometric Meaning |
|---|---|---|
| Directions are roughly the same | ||
| Two vectors perpendicular (orthogonal) | ||
| Roughly opposite directions |
Why Vectors Are Crucial for AI
In machine learning, a sample data point (an image, the embedding of a piece of text, a user's features) is essentiallya vector in high-dimensional space.。
The larger the dot product (or cosine of the angle) of the embedding vectors of two words, the more semantically similar the two words usually are. This is the mathematical basis of "similarity computation" in word vectors and recommendation systems.
What each layer of a neural network doesis essentially the dot product of the weight vector and the input vector.
Examples
import math
# Two "word vectors": assume embeddings of two words (simplified to 2D)
king = (3, 4)
queen = (4, 3)
# Norm |v| = sqrt(x² + y²)
norm_king = math.sqrt(3**2 + 4**2) # 5
norm_queen = math.sqrt(4**2 + 3**2) # 5
print(norm_king, norm_queen)
# Dot product a·b = a₁b₁ + a₂b₂
dot = king[0] * queen[0] + king[1] * queen[1]
print(dot)
# Cosine of angle cos θ = (a·b) / (|a||b|); the closer to 1, the closer the two vector directions
cos_theta = dot / (norm_king * norm_queen)
print(cos_theta)
Executing the above code outputs:
5.0 5.0 24 0.96
A cosine similarity of 0.96 means the two vector directions are very close; this is precisely the geometric expression of "semantic similarity."
Exercise: vector, find its length。
Click to see the answer
Basic Trigonometric Functions: sin and cos
From Right Triangles to the Unit Circle
When first learning trigonometric functions, one usually starts with the side-length ratios of a right triangle:
A more general and more suitable way to understand them for AI scenarios is theunit circle definition: On a circle of radius 1, from the positive x-axis rotating counterclockwise by an anglethe corresponding point has coordinates。
The advantage of this definition is:It can take any angle (not limited to 0°~90°). Sin and cos keep "going around in circles" as θ changes, exhibitingperiodicity(every full turn is 360°, i.e.,2π radians, the values repeat once).
Click "Play" to see how the point on the unit circle traces out a sine wave.
Key Properties
| Properties | Expression | Meaning |
|---|---|---|
| Boundedness | No matter how large the angle, the value always lies in [-1, 1] | |
| Periodicity | Every full turn, the value repeats once | |
| Square relation | Direct corollary of the unit circle radius being 1 (Pythagorean theorem) |
sin, cos, and Positional Encoding
When processing a sequence of text, the Transformer model inherently does not know the "order of words" (it processes all words in parallel).
To encode "position" information into the model, Transformer uses the following formula:
Here two core properties of sin/cos are used:
Boundedness: No matter how long the sequence is, the values of positional encoding will not explode; they always stay within [-1, 1] and do not interfere with other numerical computations.
Periodicity + superposition of different frequencies: Different dimensions use sin/cos waves of different frequencies, so the model can distinguish "adjacent positions" (high-frequency waves change quickly) and also perceive "long-distance positional relationships" (low-frequency waves change slowly).
Behind this is the same as in signal processingFourier transformThe idea is shared: any complex signal can be represented by a superposition of a series of sine waves of different frequencies.
Examples
import math
d = 16 # Total encoding dimension (for illustration; real model is 512+)
# The angular frequency of dimension 2i is 1 / 10000^(2i/d)
# Low dimensions have high frequency (distinguish adjacent positions), high dimensions have low frequency (perceive long distances)
for i in [0, 2, 4]:
freq = 1 / (10000 ** (i / d))
print(i, round(freq, 6))
# sin encoding value at pos=3 on dimension 0: sin(3 * 1.0)
print(round(math.sin(3 * 1.0), 6))
Executing the above code outputs:
0 1.0 2 0.316228 4 0.1 0.14112
The frequency of dimension 0 is 1.0, and dimension 4 decays to 0.1 -- the frequency decays exponentially along dimensions; this is exactly the layout of "high frequency first, low frequency later".
Coordinate Systems and Spatial Intuition
The 2D Coordinate System Is the Starting Point for Understanding 'Higher-Dimensional Spaces'
Plane rectangular coordinate systemUse two mutually perpendicular number axes to give every point on the plane a unique "address".
This intuitive idea can be extended directly: add a z-axis for three-dimensional space, and for n-dimensional space it is。
Although n-dimensional space cannot be drawn directly, the algebraic treatment is exactly the same: each coordinate axis ("dimension") is still an independent direction.
What This Means for Understanding 'Vector Spaces' in Linear Algebra
| 2D intuition | Higher-dimensional generalization | Applications in AI |
|---|---|---|
| Projection: the component of a vector along an axis | Projection of a high-dimensional vector in any direction | Dot product / orthogonal decomposition is the basis of dimensionality reduction algorithms |
| Orthogonality: two directions are perpendicular and do not interfere with each other | "Orthogonal basis" in high-dimensional space | PCA dimensionality reduction finds mutually orthogonal principal directions |
| Basis: the x-axis and y-axis are the most basic reference directions | Any vector is written as a combination of basis vectors | The dimension of the embedding space is the number of basis vectors |
The embedding dimension of many language models is from hundreds to thousands (for example, 768 dimensions). You cannot draw this space on paper, but you can "imagine" it using familiar operations from 2D and 3D.
A Concrete Analogy
Imagine you are describing a movie; you can score it on many dimensions such as "amount of action scenes," "level of humor," "plot complexity," and so on.
Each movie corresponds to a point (vector) in this "multidimensional scoring space."
The more similar two movies are, the "closer" their corresponding vectors are in space (large dot product, small angle).
This is exactly the geometric intuition behind "similar content recommendation" in recommendation systems, and the origin of this intuition is the spatial sense of "whether two points are close" in a 2D coordinate system.
Exercise: If a vector has a large projection along the x-axis direction and almost zero projection along the y-axis direction, what does this indicate about the vector's general direction?
Click to view answer
The direction is almost along the x-axis: its "contribution" in the y direction is very small, meaning it is nearly perpendicular to the y-axis and nearly parallel to the x-axis.
Chapter Summary
| Topic | One-sentence core idea |
|---|---|
| Plane vector | Carries both "direction" and "length," and is the basic unit of high-dimensional data (embedding) |
| Dot product | Measures "how close" two directions are; it is the mathematical foundation of similarity computation |
| sin / cos | Coordinates on the unit circle have periodicity and boundedness, and are a mathematical tool for positional encoding |
| Coordinate system | The spatial intuition of 2D/3D is the starting point for understanding n-dimensional vector spaces, projection, orthogonality, and other concepts |