Machine Learning - House Price Prediction

When buying a house, people usually make a comprehensive judgment: location, area, property age, number of rooms, population density, transportation conditions, etc. Every factor influences the price, but this influence is not linear or singular; rather, it is the result of accumulation, trade-offs, and competition among factors.

In machine learning,regression problems, in essence, turn this "experience-based judgment" into a computable, reusable, and evaluable mathematical model.

This chapter will start from scratch and walk through the completehouse price predictionstandard machine learning pipeline: data understanding → feature analysis → model training → model evaluation → model optimization. The goal is not to memorize APIs, but to understand what each step does and why it is done.


Part 1: Project Preparation and Environment Setup

1.1 Core Tools Used

  • NumPy: A low-level numerical computation tool that provides efficient array operations.
  • Pandas: A core tool for tabular data analysis, essential in the early stages of machine learning.
  • Matplotlib / Seaborn: Data visualization, used to understand data distributions and relationships.
  • Scikit-learn: A machine learning toolbox covering datasets, models, and evaluation methods.

1.2 Installing Dependencies

pip install numpy pandas matplotlib seaborn scikit-learn

Part 2: Loading and Understanding the Dataset

Early tutorials often used the Boston Housing dataset, but it has been deprecated. Here we use the officially recommendedCalifornia Housingdataset, which has the same concept but more standardized data.

Example: Loading Data

import pandas as pd
from sklearn.datasets import fetch_california_housing

# Load the California housing price dataset
data = fetch_california_housing()

# Feature data (X)
df_features = pd.DataFrame(
    data.data,
    columns=data.feature_names
)

# Target variable (y): median house value
df_target = pd.DataFrame(
    data.target,
    columns=['MedHouseVal']
)

print(df_features.head())
print(df_target.head())

Output:

   MedInc  HouseAge  AveRooms  AveBedrms  Population  AveOccup  Latitude  Longitude
0  8.3252      41.0  6.984127   1.023810       322.0  2.555556     37.88    -122.23
1  8.3014      21.0  6.238137   0.971880      2401.0  2.109842     37.86    -122.22
2  7.2574      52.0  8.288136   1.073446       496.0  2.802260     37.85    -122.24
3  5.6431      52.0  5.817352   1.073059       558.0  2.547945     37.85    -122.25
4  3.8462      52.0  6.281853   1.081081       565.0  2.181467     37.85    -122.25
   MedHouseVal
0        4.526
1        3.585
2        3.521
3        3.413
4        3.422

At this point, you should be clear on three things:

  • Each row is the statistical features of one house
  • Each column is a variable that can be used for prediction
  • MedHouseValis the target we want to predict

Part 3: Exploratory Data Analysis (EDA)

Before training the model, you must first answer one question:Is the data worth learning from?

3.1 Data Structure and Missing Value Check

Example: Data Overview

print("Feature dimensions:", df_features.shape)
print("Target dimensions:", df_target.shape)

print("\n"Data types and missing values:")
df_features.info()

print("\n"Missing value statistics:")
print(df_features.isnull().sum())

Conclusion:

  • Sufficient sample size (around 20,000 entries)
  • All features are numerical
  • No missing values, can directly build the model

3.2 Relationship Between a Single Feature and House Price

Machine learning is not a black box; at least at the introductory stage, you should know what the model is learning.

Example: Relationship Between Number of Rooms and House Price

import matplotlib.pyplot as plt

plt.figure(figsize=(8, 6))
plt.scatter(df_features['AveRooms'], df_target['MedHouseVal'], alpha=0.4)
plt.xlabel('Average Rooms')
plt.ylabel('Median House Value')
plt.title('Rooms vs House Price')
plt.grid(True)
plt.show()

Intuitive conclusion: the more rooms, the higher the house price overall, but there is significant dispersion. This is exactly the purpose of a regression model.

3.3 Feature Correlation Analysis

Example: Correlation Heatmap

import seaborn as sns

df_all = pd.concat([df_features, df_target], axis=1)
corr = df_all.corr()

plt.figure(figsize=(10, 8))
sns.heatmap(corr, cmap='coolwarm', center=0)
plt.title('Feature Correlation Heatmap')
plt.show()

The purpose of this step is not to "select a model," but to confirm: there is indeed a learnable statistical relationship.


Part 4: Building the First Regression Model

4.1 Splitting Training Set and Test Set

The model cannot use the same batch of data to both learn and be tested; otherwise, the evaluation results are meaningless.

Example: Dataset Splitting

from sklearn.model_selection import train_test_split

X = df_features
y = df_target['MedHouseVal']

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

print("Training set:", X_train.shape)
print("Test set:", X_test.shape)

4.2 Training a Linear Regression Model

Example: Model Training

from sklearn.linear_model import LinearRegression

model = LinearRegression()
model.fit(X_train, y_train)

print("Intercept:", model.intercept_)
print("Coefficients:")
for name, coef in zip(X.columns, model.coef_):
    print(f"{name}: {coef:.4f}")

The advantage of linear regression is:strong interpretability,suitable for understanding the essence of regression problems.


Part 5: Model Prediction and Evaluation

5.1 Comparison of Prediction Results

Example: Predictions vs. True Values

y_pred = model.predict(X_test)

result = pd.DataFrame({
    "Actual": y_test.values,
    "Predicted": y_pred
})

print(result.head())

5.2 Quantifying the Model with Evaluation Metrics

Example: Evaluation Metrics

from sklearn.metrics import mean_squared_error, r2_score

mse = mean_squared_error(y_test, y_pred)
r2 = r2_score(y_test, y_pred)

print("MSE:", mse)
print("R2:", r2)

Explanation:

  • The smaller the MSE, the lower the prediction error
  • The closer R² is to 1, the stronger the model's explanatory power

Part 6: Model Optimization – Standardization and Ridge Regression

6.1 Why Standardize?

Huge differences in feature scales will cause the model to bias toward features with larger values.

Example: Feature Standardization

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

6.2 Using Ridge Regression to Suppress Overfitting

Example: Ridge Regression

from sklearn.linear_model import Ridge

ridge = Ridge(alpha=1.0)
ridge.fit(X_train_scaled, y_train)

y_pred_ridge = ridge.predict(X_test_scaled)

print("Ridge MSE:", mean_squared_error(y_test, y_pred_ridge))
print("Ridge R2:", r2_score(y_test, y_pred_ridge))

Part 7: Full Pipeline Review

Other Extensions