Write a Multiple Linear Regression from Scratch

Comprehensively apply matrix operations + gradient descent + MSE, and fully walk through the entire process of "data → model → training → prediction".

After learning this case, you will understand:A machine learning model that can run is essentially a combination of linear algebra + calculus + optimization.


Life Introduction

Estimating House Prices — Not Just Area

House prices depend not only on area, but also on the number of bedrooms, floor, proximity to the subway, building age... Each factor has its own weight.

Multiple linear regression is:Give each factor a weight w, multiply all factors by their weights and sum them, then add a base price b.Then use historical transaction data to infer these weights — this is training.


Intuitive Understanding

The complete five-step machine learning loop:

Prepare data
X, y
Forward prediction
X*W+b
Calculate loss
MSE
Compute gradient
X.T*error
Update parameters
W-=lr*dW

These five steps are executed in a loop in each iteration until the loss converges. The fit() method in PyTorch/TensorFlow is exactly this loop internally.


Mathematical Definition

\[ \hat{y} = XW + b, \quad L = \frac{1}{n}\sum(\hat{y}_i - y_i)^2 \] \[ \frac{\partial L}{\partial W} = \frac{2}{n} X^T (XW + b - y), \quad \frac{\partial L}{\partial b} = \frac{2}{n} \sum (XW + b - y) \]

Python Hands-on Practice

Example

import numpy as np
np.random.seed(0)


class LinearRegressionFromScratch:
    """Completely implemented with NumPy matrix operations + gradient descent"""

    def __init__(self, lr=0.01, epochs=500):
        self.lr = lr
        self.epochs = epochs
        self.W = None
        self.b = None
        self.loss_history = []

    def fit(self, X, y):
        n_samples, n_features = X.shape
        self.W = np.zeros(n_features)
        self.b = 0.0

        for epoch in range(self.epochs):
            y_pred = X @ self.W + self.b
            error = y_pred - y
            loss = np.mean(error ** 2)
            self.loss_history.append(loss)

            dW = (2 / n_samples) * (X.T @ error)
            db = (2 / n_samples) * np.sum(error)

            self.W -= self.lr * dW
            self.b -= self.lr * db

            if epoch % 100 == 0:
                print(f"EXAMPLE epoch {epoch:4d}  loss={loss:.4f}")

    def predict(self, X):
        return X @ self.W + self.b


# Construct data: y = 3*x1 - 2*x2 + 5 + noise
X = np.random.uniform(-5, 5, size=(200, 2))
true_W = np.array([3.0, -2.0])
true_b = 5.0
y = X @ true_W + true_b + np.random.randn(200) * 1.5

model = LinearRegressionFromScratch(lr=0.01, epochs=500)
model.fit(X, y)

print(f"\nEXAMPLE learned: W={np.round(model.W, 3)}, b={round(model.b, 3)}")
print(f"True: W={true_W}, b={true_b}")
print(f"Loss: {model.loss_history[0]:.1f} -> {model.loss_history[-1]:.2f}")

# Predict new samples
X_new = np.array([[1.0, 1.0], [-2.0, 3.0]])
print(f"Prediction: {np.round(model.predict(X_new), 2)}")
EXAMPLE epoch    0  loss=157.0828
EXAMPLE epoch  100  loss=2.6062
EXAMPLE epoch  200  loss=2.3077
EXAMPLE epoch  300  loss=2.2581
EXAMPLE epoch  400  loss=2.2475

EXAMPLE 学到的: W=[ 3.014 -2.017], b=4.737
真实的:       W=[ 3. -2.], b=5.0
损失: 157.1 -> 2.25
预测: [5.73 -4.18]

Application Scenarios in AI

ScenarioConnection to this case
Baseline modelThe first version model for any regression task should be linear regression — simple and interpretable
The last layer of a neural networkMany networks' output layer is a linear layer — essentially linear regression
Feature importance analysisAfter training, look at the size of W — features with larger absolute values have a greater impact on the result
Other extensions