Pandas df.replace() function

Common Pandas functions


df.replace()It is a function in Pandas used to replace data in a DataFrame.

Data replacement is a common operation in data cleaning,replace()It can help you replace specific values with new values, and supports multiple flexible methods such as single-value replacement, multi-value replacement, and regular expression replacement. This is very useful in scenarios such as handling outliers, unifying data formats, and encoding conversion.


Basic syntax and parameters

replace()It is a member function of DataFrame, through the dot operator.to call.

Syntax format

DataFrame.replace(to_replace=None, value=None, inplace=False, limit=None, regex=False, method='pad')

Parameter description

Parameter Type Required Description Default value
to_replace str, regex, list, dict, int, float, None Required The value to be replaced. It can be a single value, a list of values, a dictionary, a regular expression, etc. None
value scalar, dict, list, str, regex, None Optional The value after replacement. Ifto_replaceit is a dictionary, this parameter can be omitted; if it is not a dictionary, this parameter is required. None
inplace bool Optional If it isTrue, modify the original DataFrame directly without returning a new object; if it isFalse, return a new DataFrame, the original data remains unchanged. False
limit int Optional Specify the maximum number of replacements. None
regex bool or str Optional If it isTrue, treatto_replaceas a regular expression. False
method str Optional Whento_replaceUsed when it is a list.'pad'or'ffill'Indicates forward fill;'backfill'or'bfill'Indicates backward fill. 'pad'

Return value description

  • Returns a new DataFrame (ifinplace=False), orNone(ifinplace=True)。
  • The specified values in the returned DataFrame have been replaced.

Examples

Let's thoroughly master through a series of examplesreplace()the usage of.

Example 1: Single value replacement

Replace a specific value in the DataFrame with a new value.

Example

import pandas as pd
import numpy as np

# Create a DataFrame that contains duplicate data
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'Department': ['Technology', 'Marketing', 'Technology', 'Marketing'],
    'Salary': [5000, 6000, 5500, 7000]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Replace 'technology' with 'R&D'
df_replaced = df.replace('Technology', R&D)

print("Replaced Data:")
print(df_replaced)

Expected output:

原始数据:
    姓名  部门   薪资
0  张三  技术  5000
1  李四  市场  6000
2  王五  技术  5500
3  赵六  市场  7000
==================================================
替换后的数据:
    姓名   部门  薪资
0  张三  研发  5000
1  李四  市场  6000
2  王五  研发  5500
3  赵六  市场  7000

Code analysis:

  1. There are two rows with "Technology" department in the DataFrame.
  2. Usedf.replace('技术', '研发')to replace all "technology" with "R&D".
  3. This method replaces all matching values in the DataFrame.

Example 2: One-to-one replacement of multiple values

Use a dictionary to replace multiple different values at once.

Example

import pandas as pd

# Create a DataFrame containing the data to be replaced
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'],
    Level: ['A', 'B', 'A', 'C']
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Use a dictionary to perform multiple replacements
replacements = {
    'Beijing': 'Beijing',
    'Shanghai': 'Shanghai',
    'Guangzhou': 'Guangzhou',
    'Shenzhen': 'Shenzhen',
    'A': 'Excellent',
    'B': Good,
    'C': 'Pass'
}
df_replaced = df.replace(replacements)

print("Replaced Data:")
print(df_replaced)

Expected output:

原始数据:
    姓名  城市  等级
0  张三  北京   A
1  李四  上海   B
2  王五  广州   A
3  赵六  深圳   C
==================================================
替换后的数据:
    姓名   城市  等级
0  张三  北京市  优秀
1  李四  上海市  良好
2  王五  广州市  优秀
3  赵六  深圳市  及格

Code analysis:

  • Through the dictionaryreplacements, we can specify multiple replacement rules at once.
  • This method is very efficient, avoiding multiple calls toreplace()。

Example 3: Replace values in a specific column

You can replace values only in specific columns without affecting other columns.

Example

import pandas as pd

# Create a DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'Department': ['Technology', 'Marketing', 'Technology', 'Marketing'],
    Position: ['Technology', 'Marketing', 'Technology', 'Marketing']
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Only replace "Technology" with "R&D" in the "Department" column
df_replaced = df.replace({'Department': 'Technology'}, R&D)

print(Data after replacing only the department column:)
print(df_replaced)
print("=" * 50)

You can also use a nested dictionary to apply different replacement rules to different columns.
df_replaced2 = df.replace({'Department': {'Technology': R&D, 'Marketing': Sales}, Position: 'Technology'})

print("Data after applying different replacement rules to different columns:")
print(df_replaced2)

Expected output:

原始数据:
    姓名  部门  职位
0  张三  技术  技术
1  李四  市场  市场
2  王五  技术  技术
3  赵六  市场  市场
==================================================
只替换部门列后的数据:
    姓名   部门  职位
0  张三  研发  技术
1  李四  市场  市场
2  王五  研发  技术
3  赵六  市场  市场
==================================================
对不同列使用不同规则替换后的数据:
    姓名   部门  职位
0  张三  研发  研发
1  李四  销售  市场
2  王五  研发  研发
3  赵六  销售  市场

Code analysis:

  • Using{'department': 'Technology'}can replace only the values in the "Department" column.
  • Nested dictionary{'department': {'Technology': 'R&D'}}can specify different replacement rules for different columns.

Example 4: Replace using regular expressions

regex=TrueThe parameter allows the use of regular expressions for replacement, which is very useful when handling text with pattern matching.

Example

import pandas as pd

# Create a DataFrame that needs to be processed with regular expressions
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    Phone: ['138-0000-0000', '139-1111-1111', '137-2222-2222', '136-3333-3333'],
    Remarks: [Normal, 'VIPclient', Normal, 'VIP customer']
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Use regular expressions to replace separators in phone numbers
df_replaced = df.replace(to_replace=r'(d{3})-(d{4})-(d{4})', value=r'1****3', regex=True)

print(Data after replacing phone numbers (hiding the middle 4 digits):)
print(df_replaced)
print("=" * 50)

# Replace "VIP customer" in the remark (there may be whitespace differences)
df_replaced2 = df.replace(to_replace=r'VIPs*Customers?', value='VIP', regex=True)

print(Data after unifying VIP remarks:)
print(df_replaced2)

Expected output:

原始数据:
    姓名          电话        备注
0  张三  138-0000-0000       正常
1  李四  139-1111-1111   VIP客户
2  王五  137-2222-2222       正常
3  赵六  136-3333-3333  VIP 客户
==================================================
替换电话号码(隐藏中间4位)后的数据:
    姓名          电话  备注
0  张三  138****0000   正常
1  李四  139****1111  VIP客户
2  王五  137****2222   正常
3  赵六  136****3333  VIP 客户
==================================================
统一VIP备注后的数据:
    姓名          电话  备注
0  张三  138-0000-0000  正常
1  李四  139-1111-1111  VIP
2  王五  137-2222-2222  正常
3  赵六  136-3333-3333  VIP

Code analysis:

  • The first example uses a regular expression(d{3})-(d{4})-(d{4})to match the phone number format and replace it with1****3, hiding the middle 4 digits.
  • The second example usesVIPs*客户?to match both "VIPclient" and "VIP client" formats, and uniformly replace them with "VIP".

Example 5: Replace missing values NaN

replace()It can also be used to replace missing valuesNaN, which is a common operation in data cleaning.

Example

import pandas as pd
import numpy as np

# Create a DataFrame with missing values
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
    'Age': [25, np.nan, 35, np.nan],
    'Salary': [5000, 6000, np.nan, 8000]
}
df = pd.DataFrame(data)

print("Original data:")
print(df)
print("=" * 50)

# Replace NaN with 0
df_replaced = df.replace(np.nan, 0)

print("Data after replacing NaN with 0:")
print(df_replaced)
print("=" * 50)

# Replace NaN with other values, such as "Unknown"
df_replaced2 = df.replace(np.nan, 'Unknown')

print("Data after replacing NaN with 'Unknown':")
print(df_replaced2)

Expected output:

原始数据:
    姓名   年龄    薪资
0  张三  25.0  5000.0
1  李四   NaN  6000.0
2  王五  35.0     NaN
3  赵六   NaN  8000.0
==================================================
将 NaN 替换为 0 后的数据:
    姓名   年龄  薪资
0  张三  25.0  5000
1  李四   0.0  6000
2  王五  35.0     0
3  赵六   0.0  8000
==================================================
将 NaN 替换为'未知'后的数据:
    姓名   年龄    薪资
0  张三  25.0  5000.0
1  李四  未知  6000.0
2  王五  35.0    未知
3  赵六  未知  8000.0

Code explanation:

  • np.nanrepresents missing values. You can usereplace(np.nan, value)to replace them.
  • Choose an appropriate value to replace missing values based on the data type and business requirements.

Example 6: Numeric replacement

You can replace numeric data to perform data standardization or outlier handling.

Example

import pandas as pd

# Create a DataFrame containing outliers
data = {
    'Student': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
    'Score': [85, 150, 78, 92, -10]  # 150 and -10 are outliers
}
df = pd.DataFrame(data)

print("Original data:")
print(df)
print("=" * 50)

# Replace outliers with values in a reasonable range
df_replaced = df.replace({150: 100, -10: 0})

print("Data after replacing outliers:")
print(df_replaced)
print("=" * 50)

# Use conditional logic to replace multiple values
df_replaced2 = df.replace(to_replace=[150, -10], value=[100, 0])

print("Using lists to replace multiple values:")
print(df_replaced2)

Expected output:

原始数据:
    学生  成绩
0  张三    85
1  李四   150
2  王五    78
3  赵六    92
4  钱七   -10
==================================================
将异常值替换后的数据:
    学生  成绩
0  张三    85
1  李四   100
2  王五    78
3  赵六    92
4  钱七     0
==================================================
使用列表替换多个值:
    学生  成绩
0  张三    85
1  李四   100
2  王五    78
3  赵六    92
4  钱七     0

Code explanation:

  • dictionary{150: 100, -10: 0}Replace outlier 150 with 100, and -10 with 0.
  • list[150, -10]and[100, 0]Specify the values to be replaced and the replacement values respectively, corresponding in order.

Notes

  • replace()By default, the original DataFrame is not modified. If you want to modify it in place, useinplace=Trueparameter.
  • When using regular expressions, ensure the regex syntax is correct; complex regular expressions may cause unexpected results.
  • Replacement is based on value matching and does not change the data type.
  • Pay attention to case sensitivity; "Technology" and "Technology" are different values.
  • Before data replacement, it is recommended to back up the original data for comparison and traceability.

Common Pandas Functions

Other Extensions