Images are Matrices

A grayscale image is essentially a two-dimensional NumPy array. When you transpose, flip, crop, or convolve the matrix, the image changes accordingly.

After completing this case, you will understand:The foundation of image processing is linear algebra; a matrix is just a table of numbers.


Everyday Introduction

Pixel grid of a photo

Zoom in on a photo on your phone, and zoom in again—until you can see tiny squares. Each square has only one color value: 0 means pure black, 255 means pure white, and the numbers in between are different shades of gray.

The entire photo is a rectangular array of these tiny squares—mathematicians call it a "matrix」。


Intuitive Understanding

A matrix is not a mysterious concept—it is simply a table filled with numbers.

For example, in the \(3 \times 3\) matrix below, each number represents the brightness of a pixel (0 is darkest, 1 is brightest):

3x3 grayscale image One matrix, 9 numbers
\[ M = \begin{bmatrix} 0.0 & 0.5 & 1.0 \\ 0.3 & 0.7 & 0.6 \\ 0.8 & 0.2 & 0.4 \end{bmatrix} \]

With this understanding, the various operations of image processing become clear:

Transpose
img.T
Flip along the diagonal
Horizontal mirror
img[:, ::-1]
Flip left-right
Vertical mirror
img[::-1, :]
Flip up-down
Crop
img[10:54, 10:54]
Extract subregion
Blur
Convolution kernel sliding average
Neighborhood average
Rotation
Coordinate transformation
Matrix multiplication

When you click a button in Photoshop, behind the scenes NumPy is operating on a two-dimensional array.


Mathematical Definition

A \(H \times W\) grayscale image can be represented as a real-valued matrix:

\[ M \in \mathbb{R}^{H \times W}, \quad M_{ij} \in [0, 1] \]

where \(M_{ij}\) represents the pixel brightness value at the \(i\)-th row and \(j\)-th column.

Core concepts of matrix operations:

OperationMathematical expressionNumPy syntaxImage effect
Transpose\(M^T_{ij} = M_{ji}\)img.TSwap rows and columns, flip along the diagonal
Horizontal flip\(M'_{ij} = M_{i, W-1-j}\)img[:, ::-1]Reverse column indices
Vertical flip\(M'_{ij} = M_{H-1-i, j}\)img[::-1, :]Row index reversal
Mean blur\(M'_{ij} = \frac{1}{k^2}\sum_{p,q} M_{i+p, j+q}\)Sliding window to compute mean()Replace each pixel with the neighborhood mean

The "sliding window" operation used in mean blur is essentiallyconvolutionits prototype — the core operation of Convolutional Neural Networks (CNN).


Python Hands-on Practice

Generate a synthetic grayscale image, and use NumPy array operations to implement transpose, flip, cropping, and hand-written convolution blur.

Example

import numpy as np
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
import os

# Create output directory
OUT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "outputs")
os.makedirs(OUT, exist_ok=True)

# 1. Generate a synthetic grayscale image: diagonal gradient + bright square in the middle
size = 64
img = np.zeros((size, size))
for i in range(size):
    for j in range(size):
        img[i, j] = (i + j) / (2 * size)  # Diagonal gradient from dark to bright
img[20:44, 20:44] += 0.6                  # Add a bright square in the middle
img = np.clip(img, 0, 1)                  # Clamp to [0, 1] range

print("EXAMPLE image matrix shape:, img.shape)
print("First 3x3 part of the image matrix:\n", np.round(img[:3, :3], 3))

# 2. Matrix operations = image transformations
img_transpose = img.T           # Transpose: swap rows and columns
img_flip_lr = img[:, ::-1]      # Horizontal mirror: reverse column indices
img_flip_ud = img[::-1, :]      # Vertical mirror: reverse row indices
img_crop = img[10:54, 10:54]    # Cropping: slice to get a submatrix

# 3. Hand-written mean blur (5x5 convolution kernel, without OpenCV)
def box_blur(mat, k=5):
    """Apply mean blur to the matrix using a k x k sliding window"""
    pad = k // 2
    # Pad the boundary with edge values (edge padding)
    padded = np.pad(mat, pad, mode="edge")
    out = np.zeros_like(mat)
    for i in range(mat.shape[0]):
        for j in range(mat.shape[1]):
            # Take the k x k neighborhood and compute the mean
            out[i, j] = padded[i:i+k, j:j+k].mean()
    return out

img_blur = box_blur(img, k=5)
print("EXAMPLE first 3x3 part of the blurred matrix:\n",
      np.round(img_blur[:3, :3], 3))

# 4. Save visual comparison figure
fig, axes = plt.subplots(2, 3, figsize=(12, 8))
titles = ["Original image (matrix)", "Transpose img.T",
          "Horizontal mirror img[:, ::-1]",
          "Vertical mirror img[::-1, :]",
          "Crop img[10:54, 10:54]",
          "Mean blur (5x5 convolution)"]
images = [img, img_transpose, img_flip_lr,
          img_flip_ud, img_crop, img_blur]

for ax, title, im in zip(axes.flat, titles, images):
    ax.imshow(im, cmap="gray", vmin=0, vmax=1)
    ax.set_title(title, fontsize=11)
    ax.axis("off")

plt.tight_layout()
plt.savefig(os.path.join(OUT, "01_image_as_matrix.png"), dpi=130)
print(f"\nEXAMPLE visualization result saved to {OUT}/01_image_as_matrix.png")
EXAMPLE 图像矩阵形状 (shape): (64, 64)
图像矩阵前 3x3 部分:
 [[0.    0.008 0.016]
 [0.008 0.016 0.023]
 [0.016 0.023 0.031]]
EXAMPLE 模糊后矩阵前 3x3 部分:
 [[0.009 0.011 0.013]
 [0.01  0.013 0.016]
 [0.011 0.016 0.021]]
EXAMPLE 可视化结果已保存到 .../outputs/01_image_as_matrix.png

As you can see, after blurring, the numerical differences between adjacent pixels become smaller — this is precisely the mathematical essence of "blur":Replace the original pixel value with the neighborhood mean。


Application Scenarios in AI

AI scenarioHow to use matrices
Convolutional Neural Network (CNN)The input is a 4th-order tensor composed of a batch of images (batch, H, W, C), and each layer slides a convolution kernel (small matrix) over the input
Data AugmentationDuring training, randomly flip, rotate, and crop images — all are matrix operations, used to augment the training set.
Image preprocessingBefore feeding into the model, uniformly crop to a fixed size (e.g., 224x224) — that is, matrix slicing and scaling.
Style transferRepresent the image as a feature matrix, and change the "style" features through matrix operations.

ndarray in NumPy and Tensor in PyTorch are essentially the same thing: a multidimensional numeric table. Understanding the matrix means understanding the most basic data structure in deep learning frameworks.

Other extensions