PCA for Data Dimensionality Reduction and Visualization

Implement PCA (Principal Component Analysis) by hand, without using sklearn. Intuitively see the role of eigenvalues and eigenvectors in "compressing information and preserving main differences."

After completing this case, you will understand:PCA is about finding the direction in which the data distribution is most "spread out", retaining as much information as possible with fewer dimensions.


Real-Life Introduction

Describing the shape of a fish to a friend

You don't need to precisely describe the 3D coordinates of every point on the fish. You only need to say: "about this long, this wide, this shape." With the fewest pieces of data, you capture the most critical differences.

This is exactly what PCA does: it finds the most important directions from high-dimensional data, discards unimportant details, and describes the same data with fewer numbers.


Intuitive Understanding

Suppose we have 300 3D data points. They are not uniformly distributed—most points are arranged along a certain oblique line, and only a few deviate from it.

What does this mean? It means we don't actually need coordinates in three dimensions. Two dimensions are enough—one along the oblique line direction, and one perpendicular to it. The third dimension, perpendicular to the line, has almost no variation and can be discarded.

Centering
Covariance Matrix
Eigendecomposition
Projection for Dimensionality Reduction

The larger the eigenvalue, the more "spread out" the data is in that direction—the more important that direction is.


Mathematical Definition

PCA Algorithm Steps

\[ \begin{aligned} \text{1. Centering:}&\quad X_c = X - \bar{X} \\ \text{2. Covariance matrix:}&\quad \Sigma = \frac{1}{n-1}X_c^T X_c \\ \text{3. Eigendecomposition:}&\quad \Sigma = V \Lambda V^T \\ \text{4. Select the first k eigenvectors:}&\quad V_k = [v_1, v_2, ..., v_k] \\ \text{5. Projection:}&\quad X_{\text{reduced}} = X_c \cdot V_k \end{aligned} \]

Explained Variance Ratio

\[ \text{Explained Proportion}_i = \frac{\lambda_i}{\sum_{j=1}^{d} \lambda_j} \]

where \(\lambda_i\) is the \(i\)-th largest eigenvalue. This proportion tells you how much of the total variance the \(i\)-th principal component “contributes.”


Python Hands-On Practice

Construct a 3-dimensional dataset that is actually mainly distributed along one direction, and hand-write PCA to reduce it to 2 dimensions.

Example

import numpy as np
np.random.seed(42)

# 1. Construct 3D data, mainly distributed along one direction
n = 300
base = np.random.randn(n, 1)
X = np.hstack([
    base * 3 + np.random.randn(n, 1) * 0.3,   # x1 column
    base * 1.5 + np.random.randn(n, 1) * 0.3,  # x2 column
    np.random.randn(n, 1) * 0.2,               # x3 column (noise)
])
print("EXAMPLE original data shape:", X.shape,
      "(300 samples, each sample 3-dimensional)")

# 2. Hand-write PCA (four steps)
def pca(X, n_components=2):
    # Step 1: Centering — move the data mean to 0
    X_centered = X - X.mean(axis=0)
    # Step 2: Covariance matrix — find the relationships between all pairs of variables
    cov = np.cov(X_centered, rowvar=False)
    # Step 3: Eigendecomposition — eigh is more stable for symmetric matrices
    eigvals, eigvecs = np.linalg.eigh(cov)
    # Sort eigenvalues in descending order
    order = np.argsort(eigvals)[::-1]
    eigvals = eigvals[order]
    eigvecs = eigvecs[:, order]
    # Step 4: Take the top k eigenvectors and project
    components = eigvecs[:, :n_components]
    X_reduced = X_centered @ components
    # Calculate the proportion of variance explained by each principal component
    explained_ratio = eigvals[:n_components] / eigvals.sum()
    return X_reduced, eigvals, explained_ratio

X_reduced, eigvals, explained_ratio = pca(X, n_components=2)

print("\nEXAMPLE three eigenvalues (variance size):",
      np.round(eigvals, 3))
print("Proportion of variance explained by the first 2 principal components:",
      np.round(explained_ratio, 3))
print("Cumulative retained information:",
      round(explained_ratio.sum() * 100, 1), "%")
print("Data shape after dimensionality reduction:", X_reduced.shape,
      "(Reduced from 3 dimensions to 2 dimensions)")

# 3. Key observation: the third eigenvalue is very small
print(f"\nEXAMPLE the third eigenvalue is only {eigvals[2]:.4f},"
      f"much smaller than the first two ({eigvals[0]:.2f}, {eigvals[1]:.2f})")
print("This shows the third dimension has almost no information — it can be safely discarded!")
EXAMPLE 原始数据形状: (300, 3) (300 个样本,每个样本 3 维)

EXAMPLE 三个特征值 (方差大小): [9.569 2.267 0.04 ]
前 2 个主成分解释的方差比例: [0.806 0.191]
累计保留信息: 99.7 %
降维后数据形状: (300, 2) (从 3 维降到了 2 维)

EXAMPLE 第三个特征值只有 0.0404,远小于前两个 (9.57, 2.27)
说明第三维几乎没有信息量——可以安全丢掉!

Application Scenarios in AI

AI scenarioHow to use PCA
Data visualizationReduce high-dimensional data (such as 768-dimensional word vectors) to 2 dimensions, and observe the data distribution intuitively with a scatter plot
Feature dimensionality reductionBefore training a model, remove redundant features to reduce computation and lower the risk of overfitting
Noise filteringKeep only principal components with large variance, discard those with small variance (often noise)
Image compressionApproximate face images using a combination of a few "eigenfaces" (Eigenface)
Recommender systemsSVD (the matrix version of PCA) decomposes the user-item rating matrix to discover hidden user preference patterns.
Other extensions