Pandas df.rename() function
df.rename()is a function in Pandas used to rename the column names or index of a DataFrame.
In data analysis, clear column names and indexes are very important for code readability and maintainability.rename()It allows you to flexibly modify column names and index names, making data more standardized and easier to understand. This is especially useful during data import, cleaning, and report generation.
Basic syntax and parameters
rename()is a member function of DataFrame, called via the dot operator.to invoke.
Syntax format
DataFrame.rename(mapper=None, index=None, columns=None, axis=None, inplace=False, errors='ignore')
Parameter description
| Parameter | Type | Required | Description | Default value |
|---|---|---|---|---|
| mapper | dict or function | Optional | Mapping rules for renaming an axis (row index or column names). Can be a dictionary or a function. Used with theaxisparameter. |
None |
| index | dict or function | Optional | Directly specify renaming rules for the row index. Takes precedence overmapperandaxis。 |
None |
| columns | dict or function | Optional | Directly specify renaming rules for column names. Takes precedence overmapperandaxis。 |
None |
| axis | int or str | Optional | Specifies the axis to rename.0or'index'represents the row index;1or'columns'represents column names. |
None |
| inplace | bool | Optional | If set toTrue, it modifies the original DataFrame directly and returns no new object; if set toFalse, a new DataFrame is returned, and the original data remains unchanged. |
False |
| errors | str | Optional | Controls error handling.'raise'means an exception is raised when the corresponding key is not found;'ignore'means ignoring non-existent keys. |
'ignore' |
Return value description
- Returns a new DataFrame (if
inplace=False), orNone(ifinplace=True)。 - In the returned DataFrame, the column names or index have been renamed.
Examples
Let us thoroughly master through a series of examplesrename()the usage of.
Example 1: Rename column names
Use thecolumnsparameter to conveniently rename column names.
Example
# Create a DataFrame
data = {
'name': ['Zhang San', 'Li Si', 'Wang Wu'],
'age': [25, 30, 35],
'salary': [5000, 6000, 7000]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Use a dictionary to rename column names
df_renamed = df.rename(columns={
'name': 'Name',
'age': 'age',
'salary': 'Salary'
})
print("Data after renaming column names:")
print(df_renamed)
Expected output:
原始数据:
name age salary
0 张三 25 5000
1 李四 30 6000
2 王五 35 7000
==================================================
重命名列名后的数据:
姓名 年龄 薪资
0 张三 25 5000
1 李四 30 6000
2 王五 35 7000
Code explanation:
- The original data's column names are in English, and we use a dictionary to rename them to Chinese.
- Using the
columnsparameter only affects column names, not the row index. - This is a common data normalization operation.
Example 2: Rename row index
Use theindexparameter to rename the row index.
Example
# Create a DataFrame with default index
data = {
Fruit: [Apple, Banana, Orange],
'Quantity': [10, 20, 15]
}
df = pd.DataFrame(data)
print("Original data (default index):")
print(df)
print("=" * 50)
# Use a dictionary to rename row index
df_renamed = df.rename(index={
0: 'First line',
1: 'Second line',
2: 'Third line'
})
print("Data after renaming row index:")
print(df_renamed)
Expected output:
原始数据(默认索引):
水果 数量
0 苹果 10
1 香蕉 20
2 橙子 15
==================================================
重命名行索引后的数据:
水果 数量
第一行 苹果 10
第二行 香蕉 20
第三行 橙子 15
Code explanation:
- The original data's row index is numbers like 0, 1, 2.
- Using the
indexparameter can rename them to more meaningful names. - This is especially useful when generating reports.
Example 3: Rename both column names and row index
It can rename column names and row index at the same time.
Example
# Create a DataFrame
data = {
'A': [1, 2, 3],
'B': [4, 5, 6]
}
df = pd.DataFrame(data, index=['x', 'y', 'z'])
print("Original data:")
print(df)
print("=" * 50)
# Rename both column names and row index
df_renamed = df.rename(columns={'A': Column A, 'B': Column B}, index={'x': Row X, 'y': Row Y, 'z': Row Z})
print("Data after renaming:")
print(df_renamed)
Expected output:
原始数据:
A B
x 1 4
y 2 5
z 3 6
==================================================
重命名后的数据:
列A 列B
行X 1 4
行Y 2 5
行Z 3 6
Code explanation:
- You can use both
columnsandindexparameters. - This is very convenient during data import and standardization.
Example 4: Using a function to rename
rename()It also accepts a function as a parameter, allowing batch transformation of column names or index.
Example
# Create a DataFrame with relatively complex column names
data = {
'USER_NAME': ['Zhang San', 'Li Si'],
'USER_AGE': [25, 30],
'USER_SALARY': [5000, 6000]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Use a function to convert column names to lowercase
df_lower = df.rename(columns=str.lower)
print("After converting column names to lowercase:")
print(df_lower)
print("=" * 50)
# Use a function to remove prefix (e.g., USER_)
df_clean = df.rename(columns=lambda x: x.replace('USER_', ''))
print("After removing prefix:")
print(df_clean)
print("=" * 50)
# Use a function to add prefix
df_prefix = df.rename(columns=lambda x: 'col_' + x)
print("After adding prefix:")
print(df_prefix)
Expected output:
原始数据: USER_NAME USER_AGE USER_SALARY 0 张三 25 5000 1 李四 30 6000 列名转小写后: user_name user_age user_salary 0 张三 25 数据清洗 1 李四 30 6000 ================================================== 去除前缀后: NAME AGE SALARY 0 张三 25 5000 1 李四 30 6000
Code explanation:
str.loweris the lowercase method for strings.- Using
lambdafunctions can flexibly handle complex naming rules. - This method is suitable for batch processing a large number of column names.
Example 5: Using inplace parameter for in-place modification
Useinplace=Trueto modify the original DataFrame directly without returning a new object.
Example
# Create a DataFrame
data = {
'name': ['Zhang San', 'Li Si'],
'age': [25, 30]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print(f"Original data id: {id(df)}")
print("=" * 50)
# Use inplace=True to modify in place
df.rename(columns={'name': 'Name', 'age': 'age'}, inplace=True)
print("Data after modifying with inplace=True:")
print(df)
print(f"Modified data id: {id(df)}")
Expected output:
原始数据:
name age
0 张三 25
1 李四 30
原始数据 id: 140234567890
==================================================
使用 inplace=True 修改后的数据:
姓名 年龄
0 张三 25
1 李四 30
修改后数据 id: 140234567890 # 同一个对象
Code explanation:
- Using
inplace=Trueafter, the memory address (id) of the DataFrame remains unchanged, indicating an in-place modification. - This method can save memory, especially when working with large datasets.
Example 6: Handling non-existent column names
Use theerrorsparameter to control behavior when column names do not exist.
Example
# Create a DataFrame
data = {
'A': [1, 2, 3],
'B': [4, 5, 6]
}
df = pd.DataFrame(data)
print("Original data:")
print(df)
print("=" * 50)
# Attempt to rename a non-existent column, using errors='ignore' (default)
df_renamed = df.rename(columns={'C': 'C_new'}) # Column C does not exist
print("Renaming non-existent column (C) with default behavior:")
print(df_renamed) # No change will occur
print("=" * 50)
# Attempt to rename a non-existent column, using errors='raise'
try:
df_renamed2 = df.rename(columns={'C': 'C_new'}, errors='raise')
except KeyError as e:
print(f"Exception raised: {e}")
Expected output:
原始数据: A B 0 1 4 1 2 5 2 3 6 ================================================== 重命名不存在的列(C),使用默认行为: A B 0 1 4 1 2 5 2 3 6 # 无变化 ==================================================errors ================================================== 抛出异常: 'C'
Code explanation:
- By default (
errors='ignore'), if a column does not exist, no change will occur and no error will be raised. - If set to
errors='raise', an exception will be raised.KeyErrorexception.
Example 7: Using the axis parameter
UseaxisThe parameter allows more flexible specification of the axes to rename.
Example
# Create a DataFrame
data = {
'col1': [1, 2, 3],
'col2': [4, 5, 6]
}
df = pd.DataFrame(data, index=['row1', 'row2', 'row3'])
print("Original data:")
print(df)
print("=" * 50)
# Use axis='columns' to rename columns
df_renamed_cols = df.rename({'col1': 'First column', 'col2': 'Second column'}, axis='columns')
print("After renaming columns with axis='columns':")
print(df_renamed_cols)
print("=" * 50)
# Use axis='index' to rename row index
df_renamed_idx = df.rename({'row1': 'First row', 'row2': 'Second row', 'row3': 'Third row'}, axis='index')
print("After renaming rows with axis='index':")
print(df_renamed_idx)
Expected output:
原始数据:
col1 col2
row1 1 4
row2 2 5
row3 3 6
==================================================
使用 axis='columns' 重命名列后:
第一列 第二列
row1 1 4
row2 2 5
row3 3 6
Code explanation:
axis='columns'oraxis=1Means operating on columns.axis='index'oraxis=0Means operating on row indices.- This syntax is consistent with NumPy's style.
rename()By default, the original DataFrame is not modified. To modify it in place, useinplace=Trueparameter.- Prefer using
columnsandindexparameter instead ofmapperandaxis, because the former is clearer and easier to understand. - When using a function for renaming, make sure the function returns a string type, otherwise it may cause errors.
- When working with large datasets, using
inplace=Truecan save memory. - Before renaming, it is recommended to check the current column names. You can use
df.columnsto view.
Notes
Other Extensions
Common Pandas functions