Matrix Addition and Matrix Multiplication

After understanding vector operations, level up to matrices. Matrix multiplication is the key to understanding forward propagation in neural networks — the essence of a fully connected layer is matrix multiplication.

Matrix Addition: Adding Corresponding Positions

For two matrices with the same shape, add the elements at corresponding positions — as intuitive as vector addition.

Matrix Multiplication: Dot Product of Rows and Columns

Matrix multiplication is not as simple as "multiplying corresponding positions." Why?

The deeper meaning of a matrix is a "linear transformation" — it maps one vector to another vector.

The geometric meaning of multiplying two matrices AB is: first apply transformation B, then apply transformation A. Therefore, the operation rule for AB must reflect the mathematical structure of "composite transformations," not simple element-wise multiplication.

For example: A represents "horizontal stretching by 2x," B represents "counterclockwise rotation by 90°." AB means "rotate first, then stretch" — this is completely different from BA (stretch first, then rotate), so matrix multiplication does not satisfy the commutative law.

The operation rule is:

Left matrix
i-th row →
·
Right matrix
j-th column ↓
=
Result
(i, j)

The dot product of the i-th row of the left matrix and the j-th column of the right matrix → the value at position (i, j) of the result matrix.

Dimension Matching Rules

\( (m \times \color{#e74c3c}{n}) \;\times\; (\color{#e74c3c}{n} \times p) = (m \times p) \)

The number of columns of Amust equalthe number of rows of B. The two red-highlighted n's must be equal.

The shape of the result = (number of rows of A, number of columns of B)

Matrix multiplicationdoes not satisfy the commutative law: AB ≠ BA (in general).

The dimensions may even be mismatched, making BA undefined — this is where beginners make mistakes most easily.

The essence of matrix multiplication is the composition of linear transformations

Multiplying by matrix A → complete one transformation; then multiplying by matrix B → complete the second transformation.

Composing two transformations = one (BA) transformation. This is why matrix multiplication is defined this way.


Everyday Examples

Convenience store revenue calculation

Sales matrix for three days (rows = dates, columns = products):

ColaChipsInstant noodles
Monday1053
Tuesday874
Wednesday1246

Price matrix (rows = products, columns = stores):

Store 1Store 2
Cola33.5
Chips55.5
Instant noodles44.0

Sales matrix (3×3) × price matrix (3×2) = revenue matrix (3×2).

Each cell of the result = total revenue from a certain day at a certain store.


Mathematical Definition

Matrix Addition

\[ (\mathbf{A} + \mathbf{B})_{ij} = a_{ij} + b_{ij} \]

Prerequisite: A and B must have exactly the same shape.

Matrix Multiplication

Let A be m×n, B be n×p, then C = AB is m×p:

\[ c_{ij} = \sum_{k=1}^{n} a_{ik} b_{kj} \]

Python Hands-on Practice

Examples

import numpy as np

A_example = np.array([[10, 5, 3], [8, 7, 4], [12, 4, 6]])
price = np.array([[3, 3.5], [5, 5.5], [4, 4.0]])

# Matrix addition (same shape)
bonus = np.array([[1, 0, 1], [0, 1, 0], [1, 1, 0]])
total = A_example + bonus
print("A + bonus:\n", total)

# Matrix multiplication: (3×3) @ (3×2) = (3×2)
revenue = A_example @ price
print("\nRevenue matrix A @ price:\n", revenue)
print("Shape:", revenue.shape)

# Manually verify position (0,0)
manual = A_example[0,0]*price[0,0] + A_example[0,1]*price[1,0] + A_example[0,2]*price[2,0]
print(f"\nManual verification [0,0]: {manual} == {revenue[0,0]}")

# Trap of dimension mismatch
print(f"\nA shape: {A_example.shape}")     # (3, 3)
print(f"price.T shape: {price.T.shape}")   # (2, 3)
# A @ price.T → (3,3) @ (2,3) mismatch!
# Correct: price.T @ A → (2,3) @ (3,3) = (2,3)
print(f"price.T @ A shape: {(price.T @ A_example).shape}")  # (2, 3)
收入矩阵 A @ price:
 [[67.  72. ]
 [75.  79.5]
 [80.  88. ]]
形状: (3, 2)

A 形状: (3, 3)
price.T 形状: (2, 3)
price.T @ A 形状: (2, 3)

Interactive Matrix Multiplication Visualization

Below demonstrates 2×2 matrix A multiplied by 2×1 vector v, observe how the linear transformation maps points on a circle to an ellipse:


Application Scenarios in AI

Fully Connected Layer = Matrix Multiplication

nn.Linear(d_in, d_out)The essence: weight matrix W (d_out × d_in) multiplied by input vector x (d_in dimensions), yields output (d_out dimensions).

When processing batch data, the input X is a (batch, d_in) matrix, and we compute \( Y = XW^T \) or \( Y = WX^T \). This is a single matrix multiplication completing the forward computation for the entire batch. GPUs have specialized hardware acceleration for matrix multiplication (Tensor Core).

QK^T in the Attention Mechanism

In Transformer, Q and K are both (seq_len, d_k) matrices. \( QK^T \) is seq_len×d_k multiplied by d_k×seq_len, yielding a (seq_len, seq_len) attention score matrix.

In large models at the GPT-4 level, seq_len can reach 128K, and the computational cost of this matrix multiplication accounts for a large portion of the entire inference cost. The FlashAttention algorithm specifically optimizes the memory access pattern of this matrix multiplication.

Matrix Multiplication Implementation of Convolution (im2col)

The convolution operation can be expanded into matrix multiplication: unfold the input image into a large matrix using sliding windows (im2col), then flatten the convolution kernel into a matrix, and multiply the two. This allows convolution to leverage highly optimized GEMM (General Matrix Multiplication) libraries for acceleration.

Matrix Decomposition in LoRA Fine-tuning

LoRA adds the product of two small matrices A and B alongside the pretrained weights: \( W' = W + AB \). Here AB is matrix multiplication—A is d×r, B is r×d, and the product is a d×d low-rank matrix. A single matrix multiplication achieves parameter-efficient fine-tuning.


Other extensions