Vector Magnitude and Cosine Similarity

The dot product in the previous chapter contained both length and direction information.

This chapter separates them:Magnitude alone measures size, cosine similarity alone measures directional consistency.。

Magnitude: How long is a vector?

The magnitude is the "straight-line distance" from the origin to the endpoint of the vector — a generalization of the Pythagorean theorem to higher-dimensional spaces.

L1 Norm

\( \|\mathbf{v}\|_1 = \sum |v_i| \)

Manhattan Distance: Walking Along Streets

L1 of (3,4) = 7

L2 Norm (Default)

\( \|\mathbf{v}\|_2 = \sqrt{\sum v_i^2} \)

Straight-line Distance: Flying in the Air

L2 of (3,4) = 5

Unless otherwise specified, "magnitude" or \( \|\mathbf{v}\| \) defaults to the L2 norm.

Unit Vector: Discard Length, Keep Only Direction

Any nonzero vector divided by its own magnitude yields a length-1unit vector:

\[ \hat{\mathbf{v}} = \frac{\mathbf{v}}{\|\mathbf{v}\|} \]

After normalization, the vector's direction remains unchanged, but its magnitude becomes 1.

Cosine Similarity: Only Direction, Not Magnitude

Cosine similarity = after normalizing both vectors to unit vectors, then taking the dot product:

\[ \text{cosine\_sim}(\mathbf{a}, \mathbf{b}) = \frac{\mathbf{a}\cdot\mathbf{b}}{\|\mathbf{a}\|\|\mathbf{b}\|} = \cos\theta \]

Same Direction

cos=1

Perpendicular Directions

cos=0

Opposite Directions

cos=-1

Cosine similarity only cares about direction, not length.

v = (3,4) and w = (6,8) (w = 2v), their lengths differ by a factor of two, but the cosine similarity = 1 — because they have exactly the same direction.


Real-life Examples

Magnitude: A Measure of Distance

The straight-line distance from your home to school is 5 km (L2 = 5), but walking along the streets requires a 7 km detour (L1 = 7).

Both measures have their uses; in AI, L2 is mainly used.

Cosine Similarity: Matching Music Taste

Zhang San listened to 100 songs, Li Si listened to 10 songs. With dot product, Zhang San's score is naturally higher.

But with cosine similarity, only whether the taste direction is the same matters—Zhang San (active user) and Li Si (new user) can also be compared fairly.


Mathematical Definition

Breaking Down the Core Formula

\[ \mathbf{a} \cdot \mathbf{b} = \|\mathbf{a}\| \|\mathbf{b}\| \cos\theta \]

Dot product = product of lengths × directional consistency. Conversely, we can find:

\[ \cos\theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \|\mathbf{b}\|} \]

Python Hands-on Practice

Example

import numpy as np

v_example = np.array([3, 4, 0])
w_example = np.array([6, 8, 0])  # 2 times v

# Norm
norm_v = np.linalg.norm(v_example)   # L2 default
norm_w = np.linalg.norm(w_example)
print(f"||v|| = {norm_v:.1f}, ||w|| = {norm_w:.1f}")  # 5, 10

# Normalize
unit_v = v_example / norm_v
unit_w = w_example / norm_w
print("Normalized v:", unit_v)    # [0.6, 0.8, 0.]
print("Normalized w:", unit_w)    # [0.6, 0.8, 0.] exactly the same!

# Cosine similarity
cos_sim = np.dot(v_example, w_example) / (norm_v * norm_w)
print(f"Cosine similarity = {cos_sim:.4f}")  # 1.0

# Batch compute cosine similarity matrix
embeddings = np.array([
    [0.2, 0.5, 0.1, 0.8, 0.3],
    [0.3, 0.6, 0.2, 0.7, 0.4],
    [0.9, 0.1, 0.8, 0.1, 0.9],
])
norms = np.linalg.norm(embeddings, axis=1, keepdims=True)
normalized = embeddings / norms
sim_matrix = normalized @ normalized.T
print("\n"Cosine similarity matrix:")
print(np.round(sim_matrix, 3))
||v|| = 5.0, ||w|| = 10.0
v 归一化: [0.6 0.8 0. ]
w 归一化: [0.6 0.8 0. ]
余弦相似度 = 1.0000

余弦相似度矩阵:
[[1.    0.996 0.146]
 [0.996 1.    0.182]
 [0.146 0.182 1.   ]]

Application Scenarios in AI

Semantic Search and Recommendation Systems

After encoding the query and documents into vectors, use cosine similarity to rank—to find the most relevant documents. Google's semantic search and vector retrieval in RAG (Retrieval-Augmented Generation) all rely on cosine similarity.

Compared with dot product, cosine similarity is not affected by vector norm—a long article's vector norm may be large, but cosine similarity can fairly compare directional consistency with a short query.

Face Recognition: Cosine Similarity > Euclidean Distance

Face recognition models such as FaceNet map faces to 128-dimensional embedding vectors. Whether two faces match is determined using cosine similarity—it is insensitive to lighting conditions and photo brightness (which affect vector norm).

Two photos of the same person under different lighting have embedding vectors with almost identical direction (cos ≈ 0.95+), but the norms may differ greatly.

L2 Regularization to Prevent Overfitting

Adding \( \lambda\|\mathbf{w}\|^2 \) (weight decay) to the loss function penalizes excessively large weight norms. A large weight norm means the model overtrusts certain features and generalizes poorly. L2 regularization pulls it near the origin.

Normalization in Batch Normalization

BN normalizes the activations of each layer: subtract the mean and divide by the standard deviation, turning the data into a distribution with stable norm. This uses the concept of L2 norm to control the scale of each layer's output.


Other extensions