Matrix Multiplication for Neural Network Forward Propagation
Write \(y = Wx + b\) by hand, input matrices of different shapes, and observe how the dimensions change and match.
After completing this case study, you will understand:When neural network training reports a shape mismatch error, the essence is that the matrix multiplication dimensions do not match.
Life Introduction
Express Sorting Center
A batch of packages comes in from the conveyor belt (input), is diverted by sorting machines to different trucks (hidden layer), and finally delivered to various residential communities (output).
Each sorting machine is a "neuron" — it receives all packages and decides according to rules which packages go its own way. The sorting machine's rules are the "weights": distribute more packages in one direction, and fewer in another.
The operation of the entire sorting center is:Matrix multiplication + activation function + bias。
Intuitive Understanding
One layer of neural network = three steps, done in one line of code:
\( z = X @ W + b \)
Each part has a clear role:
(batch, in)
(in, out)
(out,)
(batch, out)
The input feature count (in_features) must equal the row count of the weight matrix. The output feature count (out_features) equals the column count of the weight matrix. batch_size remains unchanged throughout the entire process.
Mathematical Definition
Single-Layer Forward Propagation
\[ z_j = \sum_{i=1}^{n} x_i \cdot W_{ij} + b_j \]Each output \(z_j\) is the dot product of the input vector and the \(j\)-th column of the weight matrix, plus the bias.
Two-Layer Network Cascade
\[ \begin{aligned} h &= \text{ReLU}(X W_1 + b_1) \quad & (m \times 4) \cdot (4 \times 6) &= (m \times 6) \\ o &= \text{Sigmoid}(h W_2 + b_2) \quad & (m \times 6) \cdot (6 \times 2) &= (m \times 2) \end{aligned} \]Note: The output feature count (6) of the first layer must equal the row count (6) of the second layer's weight matrix. This is the "dimensional connection between layers".
Python Hands-On Practice
Use a generic dense_layer function to demonstrate three scenarios: single sample, batch processing, and two-layer network.
Example
def dense_layer(X, W, b, activation=None):
"""
Forward propagation of a single fully connected layer
X: input matrix (batch_size, in_features)
W: weight matrix (in_features, out_features)
b: bias vector (out_features,)
"""
z = X @ W + b # Matrix multiplication + broadcast addition
if activation == "relu":
return np.maximum(0, z) # ReLU: negative values become 0
if activation == "sigmoid":
return 1 / (1 + np.exp(-z)) # Sigmoid: squash to (0,1)
return z
print("=" * 55)
print("EXAMPLE Scenario 1: single sample, 4 dimensions -> 3 dimensions")
print("=" * 55)
X = np.array([[1.0, 2.0, 0.5, -1.0]]) # (1, 4)
W = np.random.randn(4, 3) * 0.5 # (4, 3)
b = np.zeros(3) # (3,)
y = dense_layer(X, W, b, activation="relu")
print("X.shape:", X.shape, " W.shape:", W.shape,
" -> y.shape:", y.shape)
print("\n" + "=" * 55)
print("EXAMPLE Scenario 2: batch of 8 samples, same weights")
print("=" * 55)
X_batch = np.random.randn(8, 4) # (8, 4)
y_batch = dense_layer(X_batch, W, b, activation="relu")
print("X_batch.shape:", X_batch.shape,
" -> y_batch.shape:", y_batch.shape)
print("Batch size 8 unchanged, feature dimension 4 -> 3")
print("\n" + "=" * 55)
print("EXAMPLE Scenario 3: stacking a two-layer network")
print("=" * 55)
# First layer: 4 dimensions -> 6 dimensions + ReLU
W1 = np.random.randn(4, 6) * 0.5
b1 = np.zeros(6)
# Second layer: 6 dimensions -> 2 dimensions + Sigmoid
W2 = np.random.randn(6, 2) * 0.5
b2 = np.zeros(2)
h = dense_layer(X_batch, W1, b1, activation="relu")
out = dense_layer(h, W2, b2, activation="sigmoid")
print(f"Input layer: {X_batch.shape}")
print(f"Hidden layer: {h.shape} (4->6, W1=(4,6))")
print(f"Output layer: {out.shape} (6->2, W2=(6,2))")
# Verification: what happens if dimensions don't match?
print("\n" + "=" * 55)
print("EXAMPLE Scenario 4: dimension mismatch causes an error")
print("=" * 55)
try:
W_wrong = np.random.randn(5, 3) * 0.5 # Number of rows is incorrect
dense_layer(X_batch, W_wrong, b)
except ValueError as e:
print("Error reason: X_batch has 4 columns,"
"W_wrong has 5 rows -> cannot be multiplied")
print(f" Correct approach: the number of rows of W ({W_wrong.shape[0]})"
f" should equal the number of columns of X ({X_batch.shape[1]})")
======================================================= EXAMPLE 场景 1:单样本,4 维 -> 3 维 ======================================================= X.shape: (1, 4) W.shape: (4, 3) -> y.shape: (1, 3) ======================================================= EXAMPLE 场景 2:8 个样本批量,同权重 ======================================================= X_batch.shape: (8, 4) -> y_batch.shape: (8, 3) 批大小 8 不变,特征维度 4 -> 3 ======================================================= EXAMPLE 场景 3:堆叠两层网络 ======================================================= 输入层: (8, 4) 隐藏层: (8, 6) (4->6, W1=(4,6)) 输出层: (8, 2) (6->2, W2=(6,2)) ======================================================= EXAMPLE 场景 4:维度不匹配会报错 ======================================================= 报错原因:X_batch 有 4 列,W_wrong 有 5 行 -> 无法相乘 正确做法:W 的行数 (5) 应等于 X 的列数 (4)
Application Scenarios in AI
| AI scenario | Dimension change example | Description |
|---|---|---|
| Fully connected layer | (32, 784) @ (784, 128) = (32, 128) | MNIST classification: 28x28=784 pixels → 128-dimensional features |
| QKV of Transformer | (B, L, D) @ (D, D) = (B, L, D) | Self-attention: input and output token dimensions remain unchanged |
| Embedding lookup table | (B, L) → (B, L, D) | Word ID sequence → word vector sequence |
| Classification head | (B, 768) @ (768, 10) = (B, 10) | Classification output of BERT's last layer |