Data Bias

Machine learning is changing the world, from recommendation systems to self-driving cars; its applications are everywhere.

However, these intelligent systems are not perfect; they share a common"Achilles' Heel" (Achilles' Heel)—Data bias.

Today, we will delve into this core issue that affects the fairness and accuracy of machine learning models.


What is data bias?

Data bias refers to training data that cannot accurately represent real-world conditions, causing machine learning models to learn incorrect patterns or make biased predictions. Simply put, it is "garbage in, garbage out"—if the input data has problems, the output results will also be flawed.

Three main types of data bias

  • Selection bias:Occurs when the data collection process itself has systematic bias. For example, using social media only to survey young people's opinions about a product, while ignoring older groups who do not use social media.

  • Measurement bias:Errors occur in the data itself during measurement or recording. For example, a facial recognition system developed primarily using photos of lighter-skinned people results in lower recognition accuracy for darker-skinned people.

  • Confirmation bias:Researchers or data annotators bring their own subjective biases into the data. For example, in sentiment analysis tasks, annotators may interpret text sentiment based on their own cultural background, ignoring expressions from other cultures.


How does data bias affect machine learning models?

Model performance degradation

When a model performs well on training data but poorly on real-world data, there is likely a data bias problem.

Example

# Simulate the impact of data bias on model performance
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

np.random.seed(42)
n_samples = 1000

# Real-world data
X_real = np.random.uniform(-5, 5, n_samples).reshape(-1, 1)
y_real = (X_real.flatten() > 0).astype(int)

# Biased data: 90% from X > 0, 10% from X < 0
mask_pos = X_real.flatten() > 0
mask_neg = X_real.flatten() <= 0

X_pos = X_real[mask_pos]
y_pos = y_real[mask_pos]

X_neg = X_real[mask_neg]
y_neg = y_real[mask_neg]

# Only sample a small number of negative samples
neg_sample_size = int(0.1 * len(X_pos))
neg_indices = np.random.choice(len(X_neg), neg_sample_size, replace=False)

X_biased = np.vstack([X_pos, X_neg[neg_indices]])
y_biased = np.hstack([y_pos, y_neg[neg_indices]])

# Split training / test
X_train, X_test, y_train, y_test = train_test_split(
    X_biased, y_biased, test_size=0.2, random_state=42
)

# Train the model
model = LogisticRegression()
model.fit(X_train, y_train)

# Training set evaluation
train_pred = model.predict(X_train)
train_acc = accuracy_score(y_train, train_pred)
print(f"Training accuracy: {train_acc:.2%}")

# Real-world evaluation
real_pred = model.predict(X_real)
real_acc = accuracy_score(y_real, real_pred)
print(f"Real-world accuracy: {real_acc:.2%}")

print("Conclusion: The training set performs well, but due to distribution bias, real-world performance significantly decreases")

The output result is:

训练准确率: 99.08%
真实世界准确率: 96.00%
结论:训练集表现良好,但由于分布偏差,真实世界性能显著下降

Fairness issues

Data bias may cause models to produce discriminatory outcomes for certain groups. For example, if a hiring algorithm is trained primarily on historical data of male employees, it may be biased against female job applicants.

Insufficient generalization ability

The model cannot adapt to new, unseen data scenarios because the training data does not cover a sufficiently diverse range of situations.


Common sources of data bias

Problems in the data collection phase

Problem type Specific manifestation Example
Sampling bias The data sample cannot represent the population Collecting autonomous driving data only in urban areas, ignoring rural roads
Temporal bias Data is outdated or not time-sensitive Using 2010 e-commerce data to predict 2023 consumption trends
Survivorship bias Only focusing on "surviving" data points Only studying data from successful companies, ignoring the experiences of failed companies

Problems in the data labeling phase

Example

# Simulate the impact of labeling bias
import pandas as pd

# Create a simulated dataset
data = {
    'text': [
        'This product is really great, I like it very much!',
        'Not sure if this works well',
        'Absolutely do not buy this junk product',
        'It's okay, just average',
        'Highly recommended, great value for money'
    ],
    # Assume annotators have their own bias: positive reviews are labeled as 1, everything else is labeled as 0
    'biased_label': [1, 0, 0, 0, 1],  # Biased labeling
    'true_label': [1, 0.5, 0, 0.5, 1]  # Real continuous sentiment scores
}

df = pd.DataFrame(data)
print("Labeling bias example:")
print(df)
print("\nProblem: The annotator incorrectly labeled neutral reviews as negative!")

The output is:

标注偏差示例:
              text  biased_label  true_label
0  这个产品真的很棒,我非常喜欢!             1         1.0
1       不太确定这个好不好用             0         0.5
2      绝对不要买这个垃圾产品             0         0.0
3          还行吧,一般般             0         0.5
4        超级推荐,物超所值             1         1.0

问题:标注者将中性评价错误地标注为负面!

Problems in the data preprocessing phase

  • Improper handling of outliers
  • Biased feature selection
  • Inappropriate data standardization methods

How to detect data bias?

1. Statistical data analysis

Check the distribution of different groups in the dataset.

Example

# Check the distribution of different groups in the dataset
import matplotlib.pyplot as plt

# -------------------------- Set Chinese font start --------------------------
plt.rcParams['font.sans-serif'] = [
    # Windows priority
    'SimHei', 'Microsoft YaHei',
    # macOS priority
    'PingFang SC', 'Heiti TC',
    # Linux priority
    'WenQuanYi Micro Hei', 'DejaVu Sans'
]
# Fix the issue of negative sign displaying as a square
plt.rcParams['axes.unicode_minus'] = False
# -------------------------- Set Chinese font end --------------------------

# Simulate demographic data
groups = ['Group A', 'Group B', 'Group C', 'Group D']
population_percent = [40, 30, 20, 10]  # Real population proportions
dataset_percent = [70, 20, 8, 2]       # Proportions in the dataset

fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 5))

# Real world population distribution
ax1.pie(population_percent, labels=groups, autopct='%1.1f%%')
ax1.set_title('Real-world population distribution')

# Dataset distribution
ax2.pie(dataset_percent, labels=groups, autopct='%1.1f%%')
ax2.set_title('Group distribution in the dataset')

plt.tight_layout()
plt.show()
print("Detection result: Group A is overrepresented in the dataset, while Group D is underrepresented!")

2. Model performance difference analysis

Compare model performance across different subgroups.

3. Fairness metric calculation

Use statistical metrics to quantify the degree of fairness of the model.


Strategies to address data bias

Data-level solutions

1. Improved data collection strategies

  • Active sampling: Consciously collect underrepresented data
  • Data augmentation: Increase data diversity through technical means
  • Multiple data sources: Integrate data from different sources

2. Data preprocessing techniques

Example

# Use resampling techniques to balance the dataset
import numpy as np
from sklearn.utils import resample

np.random.seed(42)

# 1. Construct imbalanced data (one-dimensional features + labels)
X_majority = np.random.normal(0, 1, 900).reshape(-1, 1)
y_majority = np.zeros(900, dtype=int)

X_minority = np.random.normal(2, 1, 100).reshape(-1, 1)
y_minority = np.ones(100, dtype=int)

print(f"Before resampling:")
print(f"Majority class: {len(X_majority)}")
print(f"Minority class: {len(X_minority)}")

# 2. Upsample the minority class
X_minority_upsampled, y_minority_upsampled = resample(
    X_minority,
    y_minority,
    replace=True,
    n_samples=len(X_majority),
    random_state=42
)

# 3. Merge the balanced dataset
X_balanced = np.vstack([X_majority, X_minority_upsampled])
y_balanced = np.hstack([y_majority, y_minority_upsampled])

print(f"\n"After resampling:")
print(f"Majority class: {np.sum(y_balanced == 0)}")
print(f"Minority class: {np.sum(y_balanced == 1)}")

print("\n"Conclusion: The sample counts are completely balanced")

Output:

重采样前:
多数类: 900
少数类: 100

重采样后:
多数类: 900
少数类: 900

结论:样本数量已完全平衡

Algorithm-level solutions

  • Fairness constraints: Add fairness constraints during model training.

  • Adversarial debiasing: Use adversarial learning techniques to reduce bias in the model.

  • Post-processing methods: Adjust model predictions to improve fairness.


Hands-on exercise: Build an unbiased classifier

Let's use a complete example to learn how to handle data bias throughout the entire process, from data collection to model evaluation.

Example

:.2%}"
import numpy as np
import pandas as pd
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report, confusion_matrix
from imblearn.over_sampling import SMOTE
import warnings
warnings.filterwarnings("ignore", category=FutureWarning)

# 1. Generate Simulated Imbalanced Data
X, y = make_classification(
    n_samples=2000,
    n_features=10,
    n_informative=8,
    n_redundant=2,
    n_clusters_per_class=1,
    weights=[0.9, 0.1],      # Deliberately create class bias
    flip_y=0,
    random_state=42
)

# Convert to DataFrame for analysis
feature_names = [f'feature_{i}' for i in range(X.shape[1])]
df = pd.DataFrame(X, columns=feature_names)
df['target'] = y

print("=== Data Bias Analysis ===")
print(f"Dataset shape: {df.shape}")
print("Class distribution:")
print(df['target'].value_counts())
print(f"Minority class proportion: {df['target'].value_counts(normalize=True))

# 2. Data splitting (preserve class proportions to prevent secondary bias)
X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.3,
    random_state=42,
    stratify=y
)

print("\n"=== Train/Test Set Distribution ===")
print("Training set:", np.bincount(y_train))
print("Test set:", np.bincount(y_test))

# 3. Use SMOTE to handle training set bias
print("\n"=== Handling Data Bias (SMOTE) ===")
smote = SMOTE(random_state=42)
X_train_balanced, y_train_balanced = smote.fit_resample(X_train, y_train)

print("Before SMOTE:", np.bincount(y_train))
print("After SMOTE:", np.bincount(y_train_balanced))

# 4. Model training
print("\n"=== Model Training ===")

# Baseline model: no bias handling
model_imbalanced = RandomForestClassifier(
    n_estimators=200,
    random_state=42
)
model_imbalanced.fit(X_train, y_train)

# Bias-handled model: trained after SMOTE
model_balanced = RandomForestClassifier(
    n_estimators=200,
    random_state=42
)
model_balanced.fit(X_train_balanced, y_train_balanced)

# 5. Model evaluation
print("\n"=== Model Evaluation (Test Set) ===")

print("\n"【Unhandled bias】Classification report")
y_pred_imbalanced = model_imbalanced.predict(X_test)
print(classification_report(y_test, y_pred_imbalanced, digits=4))
print("Confusion matrix:")
print(confusion_matrix(y_test, y_pred_imbalanced))

print("\n"【After SMOTE processing】Classification report")
y_pred_balanced = model_balanced.predict(X_test)
print(classification_report(y_test, y_pred_balanced, digits=4))
print("Confusion matrix:")
print(confusion_matrix(y_test, y_pred_balanced))

print(
    "\nConclusion:\n"
    "1. Without bias handling, the model has extremely low Recall on the minority class\n"
    "2. SMOTE significantly improves Recall and F1 for the minority class\n"
    "3. Accuracy is not a reliable metric for imbalanced data"
)

Output results:

=== 数据偏差分析 ===
数据集形状: (2000, 11)
类别分布:
target
0    1800
1     200
Name: count, dtype: int64
少数类占比: 10.00%

=== 训练 / 测试集分布 ===
训练集: [1260  140]
测试集: [540  60]

=== 处理数据偏差(SMOTE)===
SMOTE 前: [1260  140]
SMOTE 后: [1260 1260]

=== 模型训练 ===

=== 模型评估(测试集)===

【未处理偏差】分类报告
              precision    recall  f1-score   support

           0     0.9872    1.0000    0.9936       540
           1     1.0000    0.8833    0.9381        60

    accuracy                         0.9883       600
   macro avg     0.9936    0.9417    0.9658       600
weighted avg     0.9885    0.9883    0.9880       600

混淆矩阵:
[[540   0]
 [  7  53]]

【SMOTE 处理后】分类报告
              precision    recall  f1-score   support

           0     0.9944    0.9944    0.9944       540
           1     0.9500    0.9500    0.9500        60

    accuracy                         0.9900       600
   macro avg     0.9722    0.9722    0.9722       600
weighted avg     0.9900    0.9900    0.9900       600

混淆矩阵:
[[537   3]
 [  3  57]]

结论:
1. 不处理偏差时,模型对少数类 Recall 极低
2. SMOTE 显著提升少数类 Recall 与 F1
3. Accuracy 不是不平衡数据的可靠指标

Industry best practices for data bias management

1. Establish a data governance framework

  • Develop standard processes for data collection and annotation
  • Regularly audit data quality
  • Establish data bias detection mechanisms

2. Build a diverse team

  • Ensure diversity in the data science team
  • Include domain experts and ethicists
  • Conduct regular bias awareness training

3. Transparency and interpretability

  • Record data sources and processing procedures
  • Provide explanations for model decisions
  • Disclose model performance across different groups

4. Continuous monitoring and updating

  • Regularly evaluate model performance in the real world
  • Establish feedback mechanisms to collect user reports
  • Update models promptly to adapt to changes

Summary and outlook

Data bias is an unavoidable challenge in machine learning, but through systematic detection and handling methods, we can significantly mitigate its impact. The key is to recognize:

  1. Data bias is everywhere: Almost all real-world datasets contain some form of bias
  2. Early detection is crucial: Addressing bias at the early stages of a project is the lowest cost
  3. Combine technology with processes: Technical solutions alone are not enough; they need to be combined with processes and management
  4. Continuous improvement process: Handling data bias is a continuous process, not a one-time task

As the demands for fairness and accountability in machine learning increase, effectively managing data bias will become a core skill for every data scientist and machine learning engineer. Remember, a good machine learning system not only needs high accuracy, but also high fairness and reliability.

Other extensions