Vector Dot Product -- From Projection to Multiplication

The dot product is the most frequently used vector operation in AI.

The computation of fully connected layers and the similarity matching in attention mechanisms—both are fundamentally dot products.

The dot product hastwo equivalent perspectivesunderstanding the bridge between them is the core takeaway of this chapter.

Perspective 1: Algebraic computation

Multiply the elements at corresponding positions, then sum them up.

\( \mathbf{a} \cdot \mathbf{b} = \sum a_i b_i \)

For example: (2,3) · (4,1) = 2×4 + 3×1 = 11

Perspective 2: Geometric intuition

Project one vector onto another, then multiply the two lengths.

\( \mathbf{a} \cdot \mathbf{b} = \|\mathbf{a}\| \|\mathbf{b}\| \cos\theta \)

The smaller the angle θ, the larger the dot product.

Both methods yield the same value—this is an extremely elegant equation in mathematics.

The sign of the dot product reveals directional relationships

<90°
Dot product > 0
Directions are roughly the same
=90°
Dot product = 0
Directions are perpendicular (orthogonal)
>90°
Dot product < 0
Directions are roughly opposite

Real-life examples

Shopping bill: quantity × unit price

You bought 2 bottles of cola (3 yuan each), 3 bags of chips (5 yuan each), and 1 bucket of instant noodles (4 yuan per bucket).

The dot product of the quantity vector (2, 3, 1) and the price vector (3, 5, 4) is the total price: 2×3 + 3×5 + 1×4 = 25 yuan.

Pushing a shopping cart: force × displacement

The force you apply to the shopping cart is a vector (with magnitude and direction), and the cart's displacement is another vector.

In physics, work = component of force in the direction of displacement × magnitude of displacement = the dot product of the force vector and the displacement vector.

If you push sideways (perpendicular direction), you do no work at all—the dot product is zero.


Mathematical definition

Algebraic definition

\[ \mathbf{a} \cdot \mathbf{b} = \sum_{i=1}^{n} a_i b_i = a_1 b_1 + a_2 b_2 + \cdots + a_n b_n \]

Geometric definition

\[ \mathbf{a} \cdot \mathbf{b} = \|\mathbf{a}\| \|\mathbf{b}\| \cos\theta \]

Bridge between the two definitions: finding the angle

\[ \cos\theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \|\mathbf{b}\|} \]

This means:Given the coordinates of two vectors, you can compute the angle between them without a protractor。

Properties of operations

PropertyFormulaExplanation
Commutative law\( \mathbf{a} \cdot \mathbf{b} = \mathbf{b} \cdot \mathbf{a} \)Order does not matter
Distributive law\( \mathbf{a} \cdot (\mathbf{b} + \mathbf{c}) = \mathbf{a}\cdot\mathbf{b} + \mathbf{a}\cdot\mathbf{c} \)You can add first, then take the dot product
Dot product with itself\( \mathbf{a} \cdot \mathbf{a} = \|\mathbf{a}\|^2 \)The dot product of a vector with itself equals the square of its length

Hands-on with Python

Example

import numpy as np

a_example = np.array([2, 3, 1])
b_example = np.array([4, 1, 2])

# Method 1: np.dot()
dot1 = np.dot(a_example, b_example)
# Method 2: @ operator (recommended)
dot2 = a_example @ b_example
# Method 3: manual verification
dot3 = sum(a_example[i] * b_example[i] for i in range(len(a_example)))

print("a · b =", dot1)   # 13
print("The three methods give the same result:", dot1 == dot2 == dot3)

# Use the geometric definition to infer the angle
norm_a = np.linalg.norm(a_example)
norm_b = np.linalg.norm(b_example)
cos_theta = dot1 / (norm_a * norm_b)
theta_deg = np.degrees(np.arccos(cos_theta))
print(f"||a||={norm_a:.2f}, ||b||={norm_b:.2f}, cosθ={cos_theta:.4f}")
print(f"Angle = {theta_deg:.1f}°")

# Verify the geometric formula
verify = norm_a * norm_b * cos_theta
print(f"||a||·||b||·cosθ = {verify:.4f} (consistent with the dot product)")

# Demonstration of different angles
print("\n[1,0] · [0.8,0.6] =", np.array([1,0]) @ np.array([0.8,0.6]), " (acute angle, positive)")
print("[1,0] · [0,1] =", np.array([1,0]) @ np.array([0,1]), " (right angle, zero)")
print("[1,0] · [-1,0] =", np.array([1,0]) @ np.array([-1,0]), " (obtuse angle, negative)")
a · b = 13
三种方法结果一致: True
||a||=3.74, ||b||=4.58, cosθ=0.7580
夹角 = 40.7°

[1,0] · [0.8,0.6] = 0.8  (锐角,正)
[1,0] · [0,1] = 0  (直角,零)
[1,0] · [-1,0] = -1.0  (钝角,负)

Interactive dot product demo

Drag the endpoints of vectors a and b to see how the dot product value and the angle change. The sign of the dot product determines acute/obtuse/right angle:


Application scenarios in AI

Fully connected layer = batch computation of dot products

A neuron computes \( y = w_1x_1 + w_2x_2 + \cdots + w_nx_n + b \). The summation part is the weight vector w and the input vector x'sdot product。

When you see a fully connected layernn.Linear(784, 256), what it does mathematically is: 256 neurons, each neuron takes its own weight vector (784-dimensional) and computes the dot product with the input vector, then adds the bias. 256 dot products computed in parallel = matrix multiplication.

Attention mechanism = dot product + Softmax

The core operation of Transformer: the Query vector and the Key vector take a dot product; the larger the result, the more 'related' the two are.

\[ \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V \]

\( QK^T \) is the multiplication of a seq_len × d_k matrix and a d_k × seq_len matrix; the value at position (i, j) in the result matrix = the i-th Query and the j-th Key'sdot productThe larger this value, the higher the "attention" of the i-th position to the j-th position.

Dividing by \( \sqrt{d_k} \) is to prevent the dot product values from becoming too large, which would cause the Softmax gradient to vanish. All modern large language models such as GPT, BERT, and Claude are based on this dot-product attention mechanism.

Similarity computation in recommendation systems

The dot product of the user vector and the item vector gives the user's "preference score" for that item. This is the core of the matrix factorization recommendation algorithm—the rating matrix ≈ user latent vector matrix × transpose of the item latent vector matrix.


Other extensions