PCA for Data Dimensionality Reduction and Visualization
Implement PCA (Principal Component Analysis) by hand, without using sklearn. Intuitively see the role of eigenvalues and eigenvectors in "compressing information and preserving main differences."
After completing this case, you will understand:PCA is about finding the direction in which the data distribution is most "spread out", retaining as much information as possible with fewer dimensions.
Real-Life Introduction
Describing the shape of a fish to a friend
You don't need to precisely describe the 3D coordinates of every point on the fish. You only need to say: "about this long, this wide, this shape." With the fewest pieces of data, you capture the most critical differences.
This is exactly what PCA does: it finds the most important directions from high-dimensional data, discards unimportant details, and describes the same data with fewer numbers.
Intuitive Understanding
Suppose we have 300 3D data points. They are not uniformly distributed—most points are arranged along a certain oblique line, and only a few deviate from it.
What does this mean? It means we don't actually need coordinates in three dimensions. Two dimensions are enough—one along the oblique line direction, and one perpendicular to it. The third dimension, perpendicular to the line, has almost no variation and can be discarded.
The larger the eigenvalue, the more "spread out" the data is in that direction—the more important that direction is.
Mathematical Definition
PCA Algorithm Steps
\[ \begin{aligned} \text{1. Centering:}&\quad X_c = X - \bar{X} \\ \text{2. Covariance matrix:}&\quad \Sigma = \frac{1}{n-1}X_c^T X_c \\ \text{3. Eigendecomposition:}&\quad \Sigma = V \Lambda V^T \\ \text{4. Select the first k eigenvectors:}&\quad V_k = [v_1, v_2, ..., v_k] \\ \text{5. Projection:}&\quad X_{\text{reduced}} = X_c \cdot V_k \end{aligned} \]Explained Variance Ratio
\[ \text{Explained Proportion}_i = \frac{\lambda_i}{\sum_{j=1}^{d} \lambda_j} \]where \(\lambda_i\) is the \(i\)-th largest eigenvalue. This proportion tells you how much of the total variance the \(i\)-th principal component “contributes.”
Python Hands-On Practice
Construct a 3-dimensional dataset that is actually mainly distributed along one direction, and hand-write PCA to reduce it to 2 dimensions.
Example
np.random.seed(42)
# 1. Construct 3D data, mainly distributed along one direction
n = 300
base = np.random.randn(n, 1)
X = np.hstack([
base * 3 + np.random.randn(n, 1) * 0.3, # x1 column
base * 1.5 + np.random.randn(n, 1) * 0.3, # x2 column
np.random.randn(n, 1) * 0.2, # x3 column (noise)
])
print("EXAMPLE original data shape:", X.shape,
"(300 samples, each sample 3-dimensional)")
# 2. Hand-write PCA (four steps)
def pca(X, n_components=2):
# Step 1: Centering — move the data mean to 0
X_centered = X - X.mean(axis=0)
# Step 2: Covariance matrix — find the relationships between all pairs of variables
cov = np.cov(X_centered, rowvar=False)
# Step 3: Eigendecomposition — eigh is more stable for symmetric matrices
eigvals, eigvecs = np.linalg.eigh(cov)
# Sort eigenvalues in descending order
order = np.argsort(eigvals)[::-1]
eigvals = eigvals[order]
eigvecs = eigvecs[:, order]
# Step 4: Take the top k eigenvectors and project
components = eigvecs[:, :n_components]
X_reduced = X_centered @ components
# Calculate the proportion of variance explained by each principal component
explained_ratio = eigvals[:n_components] / eigvals.sum()
return X_reduced, eigvals, explained_ratio
X_reduced, eigvals, explained_ratio = pca(X, n_components=2)
print("\nEXAMPLE three eigenvalues (variance size):",
np.round(eigvals, 3))
print("Proportion of variance explained by the first 2 principal components:",
np.round(explained_ratio, 3))
print("Cumulative retained information:",
round(explained_ratio.sum() * 100, 1), "%")
print("Data shape after dimensionality reduction:", X_reduced.shape,
"(Reduced from 3 dimensions to 2 dimensions)")
# 3. Key observation: the third eigenvalue is very small
print(f"\nEXAMPLE the third eigenvalue is only {eigvals[2]:.4f},"
f"much smaller than the first two ({eigvals[0]:.2f}, {eigvals[1]:.2f})")
print("This shows the third dimension has almost no information — it can be safely discarded!")
EXAMPLE 原始数据形状: (300, 3) (300 个样本,每个样本 3 维) EXAMPLE 三个特征值 (方差大小): [9.569 2.267 0.04 ] 前 2 个主成分解释的方差比例: [0.806 0.191] 累计保留信息: 99.7 % 降维后数据形状: (300, 2) (从 3 维降到了 2 维) EXAMPLE 第三个特征值只有 0.0404,远小于前两个 (9.57, 2.27) 说明第三维几乎没有信息量——可以安全丢掉!
Application Scenarios in AI
| AI scenario | How to use PCA |
|---|---|
| Data visualization | Reduce high-dimensional data (such as 768-dimensional word vectors) to 2 dimensions, and observe the data distribution intuitively with a scatter plot |
| Feature dimensionality reduction | Before training a model, remove redundant features to reduce computation and lower the risk of overfitting |
| Noise filtering | Keep only principal components with large variance, discard those with small variance (often noise) |
| Image compression | Approximate face images using a combination of a few "eigenfaces" (Eigenface) |
| Recommender systems | SVD (the matrix version of PCA) decomposes the user-item rating matrix to discover hidden user preference patterns. |