Images are Matrices
A grayscale image is essentially a two-dimensional NumPy array. When you transpose, flip, crop, or convolve the matrix, the image changes accordingly.
After completing this case, you will understand:The foundation of image processing is linear algebra; a matrix is just a table of numbers.
Everyday Introduction
Pixel grid of a photo
Zoom in on a photo on your phone, and zoom in again—until you can see tiny squares. Each square has only one color value: 0 means pure black, 255 means pure white, and the numbers in between are different shades of gray.
The entire photo is a rectangular array of these tiny squares—mathematicians call it a "matrix」。
Intuitive Understanding
A matrix is not a mysterious concept—it is simply a table filled with numbers.
For example, in the \(3 \times 3\) matrix below, each number represents the brightness of a pixel (0 is darkest, 1 is brightest):
With this understanding, the various operations of image processing become clear:
When you click a button in Photoshop, behind the scenes NumPy is operating on a two-dimensional array.
Mathematical Definition
A \(H \times W\) grayscale image can be represented as a real-valued matrix:
\[ M \in \mathbb{R}^{H \times W}, \quad M_{ij} \in [0, 1] \]where \(M_{ij}\) represents the pixel brightness value at the \(i\)-th row and \(j\)-th column.
Core concepts of matrix operations:
| Operation | Mathematical expression | NumPy syntax | Image effect |
|---|---|---|---|
| Transpose | \(M^T_{ij} = M_{ji}\) | img.T | Swap rows and columns, flip along the diagonal |
| Horizontal flip | \(M'_{ij} = M_{i, W-1-j}\) | img[:, ::-1] | Reverse column indices |
| Vertical flip | \(M'_{ij} = M_{H-1-i, j}\) | img[::-1, :] | Row index reversal |
| Mean blur | \(M'_{ij} = \frac{1}{k^2}\sum_{p,q} M_{i+p, j+q}\) | Sliding window to compute mean() | Replace each pixel with the neighborhood mean |
The "sliding window" operation used in mean blur is essentiallyconvolutionits prototype — the core operation of Convolutional Neural Networks (CNN).
Python Hands-on Practice
Generate a synthetic grayscale image, and use NumPy array operations to implement transpose, flip, cropping, and hand-written convolution blur.
Example
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import os
# Create output directory
OUT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "outputs")
os.makedirs(OUT, exist_ok=True)
# 1. Generate a synthetic grayscale image: diagonal gradient + bright square in the middle
size = 64
img = np.zeros((size, size))
for i in range(size):
for j in range(size):
img[i, j] = (i + j) / (2 * size) # Diagonal gradient from dark to bright
img[20:44, 20:44] += 0.6 # Add a bright square in the middle
img = np.clip(img, 0, 1) # Clamp to [0, 1] range
print("EXAMPLE image matrix shape:, img.shape)
print("First 3x3 part of the image matrix:\n", np.round(img[:3, :3], 3))
# 2. Matrix operations = image transformations
img_transpose = img.T # Transpose: swap rows and columns
img_flip_lr = img[:, ::-1] # Horizontal mirror: reverse column indices
img_flip_ud = img[::-1, :] # Vertical mirror: reverse row indices
img_crop = img[10:54, 10:54] # Cropping: slice to get a submatrix
# 3. Hand-written mean blur (5x5 convolution kernel, without OpenCV)
def box_blur(mat, k=5):
"""Apply mean blur to the matrix using a k x k sliding window"""
pad = k // 2
# Pad the boundary with edge values (edge padding)
padded = np.pad(mat, pad, mode="edge")
out = np.zeros_like(mat)
for i in range(mat.shape[0]):
for j in range(mat.shape[1]):
# Take the k x k neighborhood and compute the mean
out[i, j] = padded[i:i+k, j:j+k].mean()
return out
img_blur = box_blur(img, k=5)
print("EXAMPLE first 3x3 part of the blurred matrix:\n",
np.round(img_blur[:3, :3], 3))
# 4. Save visual comparison figure
fig, axes = plt.subplots(2, 3, figsize=(12, 8))
titles = ["Original image (matrix)", "Transpose img.T",
"Horizontal mirror img[:, ::-1]",
"Vertical mirror img[::-1, :]",
"Crop img[10:54, 10:54]",
"Mean blur (5x5 convolution)"]
images = [img, img_transpose, img_flip_lr,
img_flip_ud, img_crop, img_blur]
for ax, title, im in zip(axes.flat, titles, images):
ax.imshow(im, cmap="gray", vmin=0, vmax=1)
ax.set_title(title, fontsize=11)
ax.axis("off")
plt.tight_layout()
plt.savefig(os.path.join(OUT, "01_image_as_matrix.png"), dpi=130)
print(f"\nEXAMPLE visualization result saved to {OUT}/01_image_as_matrix.png")
EXAMPLE 图像矩阵形状 (shape): (64, 64) 图像矩阵前 3x3 部分: [[0. 0.008 0.016] [0.008 0.016 0.023] [0.016 0.023 0.031]] EXAMPLE 模糊后矩阵前 3x3 部分: [[0.009 0.011 0.013] [0.01 0.013 0.016] [0.011 0.016 0.021]] EXAMPLE 可视化结果已保存到 .../outputs/01_image_as_matrix.png
As you can see, after blurring, the numerical differences between adjacent pixels become smaller — this is precisely the mathematical essence of "blur":Replace the original pixel value with the neighborhood mean。
Application Scenarios in AI
| AI scenario | How to use matrices |
|---|---|
| Convolutional Neural Network (CNN) | The input is a 4th-order tensor composed of a batch of images (batch, H, W, C), and each layer slides a convolution kernel (small matrix) over the input |
| Data Augmentation | During training, randomly flip, rotate, and crop images — all are matrix operations, used to augment the training set. |
| Image preprocessing | Before feeding into the model, uniformly crop to a fixed size (e.g., 224x224) — that is, matrix slicing and scaling. |
| Style transfer | Represent the image as a feature matrix, and change the "style" features through matrix operations. |
Other extensionsndarray in NumPy and Tensor in PyTorch are essentially the same thing: a multidimensional numeric table. Understanding the matrix means understanding the most basic data structure in deep learning frameworks.