Write a Multiple Linear Regression from Scratch
Comprehensively apply matrix operations + gradient descent + MSE, and fully walk through the entire process of "data → model → training → prediction".
After learning this case, you will understand:A machine learning model that can run is essentially a combination of linear algebra + calculus + optimization.
Life Introduction
Estimating House Prices — Not Just Area
House prices depend not only on area, but also on the number of bedrooms, floor, proximity to the subway, building age... Each factor has its own weight.
Multiple linear regression is:Give each factor a weight w, multiply all factors by their weights and sum them, then add a base price b.Then use historical transaction data to infer these weights — this is training.
Intuitive Understanding
The complete five-step machine learning loop:
X, y
X*W+b
MSE
X.T*error
W-=lr*dW
These five steps are executed in a loop in each iteration until the loss converges. The fit() method in PyTorch/TensorFlow is exactly this loop internally.
Mathematical Definition
\[ \hat{y} = XW + b, \quad L = \frac{1}{n}\sum(\hat{y}_i - y_i)^2 \] \[ \frac{\partial L}{\partial W} = \frac{2}{n} X^T (XW + b - y), \quad \frac{\partial L}{\partial b} = \frac{2}{n} \sum (XW + b - y) \]Python Hands-on Practice
Example
np.random.seed(0)
class LinearRegressionFromScratch:
"""Completely implemented with NumPy matrix operations + gradient descent"""
def __init__(self, lr=0.01, epochs=500):
self.lr = lr
self.epochs = epochs
self.W = None
self.b = None
self.loss_history = []
def fit(self, X, y):
n_samples, n_features = X.shape
self.W = np.zeros(n_features)
self.b = 0.0
for epoch in range(self.epochs):
y_pred = X @ self.W + self.b
error = y_pred - y
loss = np.mean(error ** 2)
self.loss_history.append(loss)
dW = (2 / n_samples) * (X.T @ error)
db = (2 / n_samples) * np.sum(error)
self.W -= self.lr * dW
self.b -= self.lr * db
if epoch % 100 == 0:
print(f"EXAMPLE epoch {epoch:4d} loss={loss:.4f}")
def predict(self, X):
return X @ self.W + self.b
# Construct data: y = 3*x1 - 2*x2 + 5 + noise
X = np.random.uniform(-5, 5, size=(200, 2))
true_W = np.array([3.0, -2.0])
true_b = 5.0
y = X @ true_W + true_b + np.random.randn(200) * 1.5
model = LinearRegressionFromScratch(lr=0.01, epochs=500)
model.fit(X, y)
print(f"\nEXAMPLE learned: W={np.round(model.W, 3)}, b={round(model.b, 3)}")
print(f"True: W={true_W}, b={true_b}")
print(f"Loss: {model.loss_history[0]:.1f} -> {model.loss_history[-1]:.2f}")
# Predict new samples
X_new = np.array([[1.0, 1.0], [-2.0, 3.0]])
print(f"Prediction: {np.round(model.predict(X_new), 2)}")
EXAMPLE epoch 0 loss=157.0828 EXAMPLE epoch 100 loss=2.6062 EXAMPLE epoch 200 loss=2.3077 EXAMPLE epoch 300 loss=2.2581 EXAMPLE epoch 400 loss=2.2475 EXAMPLE 学到的: W=[ 3.014 -2.017], b=4.737 真实的: W=[ 3. -2.], b=5.0 损失: 157.1 -> 2.25 预测: [5.73 -4.18]
Application Scenarios in AI
| Scenario | Connection to this case |
|---|---|
| Baseline model | The first version model for any regression task should be linear regression — simple and interpretable |
| The last layer of a neural network | Many networks' output layer is a linear layer — essentially linear regression |
| Feature importance analysis | After training, look at the size of W — features with larger absolute values have a greater impact on the result |