Pandas df.reset_index() Function
df.reset_index()is a function in Pandas used to reset the row index of a DataFrame.
During data processing, the index may become discontinuous or fail to meet your needs.reset_index()It can reset the index to the default integer index (0, 1, 2, ...), or convert the current index into a column. This is very useful for data cleaning, sorting, and organizing data after filtering.
Basic Syntax and Parameters
reset_index()is a member function of DataFrame, invoked via the dot operator.to call.
Syntax Format
DataFrame.reset_index(level=None, drop=False, inplace=False, col_level=0, col_fill='')
Parameter Description
| Parameter | Type | Required | Description | Default Value |
|---|---|---|---|---|
| level | int, str, tuple, or list | Optional | Specifies the index level(s) to reset. For a MultiIndex, you can choose to reset some levels. By default, all levels are reset. | None |
| drop | bool | Optional | If it isTrue, discard the original index without converting it to a column; if it isFalse, keep the original index as a new column. |
False |
| inplace | bool | Optional | If it isTrue, modify the original DataFrame directly and do not return a new object; if it isFalse, return a new DataFrame and keep the original data unchanged. |
False |
| col_level | int or str | Optional | If the column names are a MultiIndex, specify which level to insert the index into. | 0 |
| col_fill | str | Optional | If the column names are a MultiIndex, this is used to fill in the column names of other levels when the original index is converted to a column. | '' |
Return Value Description
- Returns a new DataFrame (if
inplace=False), orNone(ifinplace=True)。 - In the returned DataFrame, the index has been reset. The original index (if
drop=False) will be kept as a new column.
Examples
Let's thoroughly master, through a series of examples,reset_index()the usage of.
Example 1: Basic Usage - Reset the Index
The most common usage is to reset the index to the default integer index.
Example
import numpy as np
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Score': [85, 92, 78]
}
df = pd.DataFrame(data)
print(Original data (default index):)
print(df)
print("=" * 50)
# Sort the data, resulting in a non-contiguous index
df_sorted = df.sort_values(by='Score', ascending=False)
print(Sorted data (non-contiguous indices):)
print(df_sorted)
print("=" * 50)
# Reset index
df_reset = df_sorted.reset_index()
print(After resetting index:)
print(df_reset)
Expected output:
原始数据(默认索引): 姓名 成绩 0 张三 85 1 李四 92 2 王五 78 ================================================== 排序后的数据(索引不连续): 姓名 成绩 1 李四 92 0 张三 85 2 王五 78 ================================================== 重置索引后: 姓名 成绩 index 0 李四 92 1 # 原来的索引作为新列 1 张三 85 0 2 王五 78 2
Code explanation:
- After sorting, the index becomes 1, 0, 2, no longer continuous 0, 1, 2.
- After using
reset_index(), the index is reset to 0, 1, 2. - The original index is kept as a new column, "index".
Example 2: Use the drop Parameter to Discard the Original Index
If you do not need to keep the original index, you can usedrop=Trueparameter.
Example
import numpy as np
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Score': [85, 92, 78]
}
df = pd.DataFrame(data)
# Sort the Data
df_sorted = df.sort_values(by='Score', ascending=False)
print("Sorted Data:")
print(df_sorted)
print("=" * 50)
# Reset index, drop the original index
df_reset = df_sorted.reset_index(drop=True)
print(After resetting index (dropping original index):)
print(df_reset)
Expected output:
排序后的数据: 姓名 成绩 1 李四 92 0 张三 85 2 王五 78 ================================================== 重置索引(丢弃原索引)后: 姓名 成绩 0 李四 92 1 张三 85 2 王五 78
Code explanation:
- After using
drop=True, the original index is directly discarded and will not be kept as a new column. - This is useful when you do not need to keep the original index information.
Example 3: Reset the Index After Filtering Data
After data filtering, the index may be discontinuous. Usereset_index()to tidy up the data.
Example
import numpy as np
# Create a DataFrame containing all student grades
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
'Score': [85, 92, 78, 90, 88]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Filter students with scores greater than 85
df_filtered = df[df['Score'] > 85]
print(Filtered data (non-contiguous indices):)
print(df_filtered)
print("=" * 50)
# Reset index
df_reset = df_filtered.reset_index(drop=True)
print(After resetting index:)
print(df_reset)
Expected output:
原始数据: 姓名 成绩 0 张三 85 1 李四 92 2 王五 78 3 赵六 90 4 钱七 88 ================================================== 筛选后的数据(索引不连续): 姓名 成绩 1 李四 92 3 赵六 90 4 钱七 88
Code explanation:
- After filtering, the retained indices are the original 1, 3, 4, not the continuous 0, 1, 2.
- Using
reset_index(drop=True)can give a continuous index, which is more convenient for subsequent processing.
Example 4: Reset the Index After Removing Duplicate Rows
After deleting duplicate data, you can also usereset_index()to tidy up the index.
Example
import numpy as np
# Create a DataFrame with duplicate rows
data = {
'Name': ['Zhang San', 'Li Si', 'Zhang San', 'Wang Wu'],
'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou']
}
df = pd.DataFrame(data)
print(Original data (including duplicate rows):)
print(df)
print("=" * 50)
# Deleterepeatline
df_unique = df.drop_duplicates()
print(After removing duplicate rows (non-contiguous indices):)
print(df_unique)
print("=" * 50)
# Reset index
df_reset = df_unique.reset_index(drop=True)
print(After resetting index:)
print(df_reset)
Expected output:
原始数据(包含重复行): 姓名 城市 0 张三 北京 1 李四 上海 2 张三 北京 # 重复 3 王五 广州 ================================================== 删除重复行后(索引不连续): 姓名 城市 0 张三 北京 1 李四 上海 3 王五 广州
Code explanation:
- After deleting duplicate rows, the retained indices are 0, 1, 3, not continuous 0, 1, 2.
- After resetting the index, the data is tidier.
Example 5: Reset the Index After Handling Missing Values
After deleting rows with missing values, the index can also be reset.
Example
import numpy as np
# Create a DataFrame with missing values
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'Score': [85, np.nan, 92, 78]
}
df = pd.DataFrame(data)
print(“Raw data:”)
print(df)
print("=" * 50)
# Delete rows with missing values
df_cleaned = df.dropna()
print(After removing missing values (non-contiguous indices):)
print(df_cleaned)
print("=" * 50)
# Reset index
df_reset = df_cleaned.reset_index(drop=True)
print(After resetting index:)
print(df_reset)
Expected output:
原始数据: 姓名 成绩 0 张三 85.0 1 李四 NaN # 缺失值 2 王五 92.0 3 赵六 78.0 ================================================== 删除缺失值后(索引不连续): 姓名 成绩 0 张三 85.0 2 王五 92.0 3 赵六 78.0
Code explanation:
- After deleting row 1 (the score of Li Si is NaN), the retained indices are 0, 2, 3.
- Using
reset_index()can give a continuous index of 0, 1, 2.
Example 6: Use the inplace Parameter to Modify In Place
Usinginplace=Trueyou can directly modify the original DataFrame.
Example
import numpy as np
# Create a DataFrame
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Score': [85, 92, 78]
}
df = pd.DataFrame(data)
# Filter data
df_filtered = df[df['Score'] > 80]
print(Filtered data (index not reset):)
print(df_filtered)
print(fData id: {id(df_filtered)})
print("=" * 50)
# Use inplace=True to reset the index in place
df_filtered.reset_index(drop=True, inplace=True)
print(After resetting with inplace=True:)
print(df_filtered)
print(fData id: {id(df_filtered)})
Expected output:
筛选后的数据(未重置): 姓名 成绩 1 李四 92 0 张三 85 数据 id: 140234567890 ================================================== 使用 inplace=True 重置后: 姓名 成绩 0 李四 92 1 张三 85 数据 id: 140234567890 # 同一个对象
Code explanation:
- After using
inplace=True, the memory address (id) of the DataFrame remains unchanged. - This method can save memory.
Example 7: Resetting a Multi-level Index
For a MultiIndex, you can reset specific levels or all levels.
Example
import numpy as np
# Create a DataFrame with a MultiIndex
df = pd.DataFrame({
'Score': [85, 92, 78, 88, 90, 95]
}, index=pd.MultiIndex.from_tuples([
('Class 1', 'Zhang San'), ('Class 1', 'Li Si'), ('Class 1', 'Wang Wu'),
('Class 2', 'Zhao Liu'), ('Class 2', 'Qian Qi'), ('Class 2', 'Sun Ba')
], names=['Class', 'Name']))
print(Original data (MultiIndex):)
print(df)
print("=" * 50)
# Reset all index levels
df_reset_all = df.reset_index()
print(Reset all index levels:)
print(df_reset_all)
print("=" * 50)
# Only reset the first index level (Class)
df_reset_first = df.reset_index(level=0)
print(Only reset the Class index:)
print(df_reset_first)
Expected output:
原始数据(多层索引):
成绩
班级 姓名
一班 张三 85
李四 92
王五 78
二班 赵六 88
钱七 90
孙八 95
==================================================
重置所有索引级别:
班级 姓名 成绩
0 一班 张三 85
1 一班 李四 92
2 一班 王五 78
3 二班 赵六 88
4 二班 钱七 90
5 二班 钱七 95
Code explanation:
- A DataFrame with a MultiIndex can use
reset_index()to convert the index to columns. - Use the
levelparameter to choose which level of the index to reset. reset_index()The original DataFrame is not modified by default. To modify it in place, use theinplace=Trueparameter.- By default (
drop=False), the original index is kept as a new column. If you do not need to keep it, usedrop=True。 - Before using
reset_index(), ensure that there is no important index information to retain. - For time series data, you may need to reset the time index after resetting the index.
- In chained operations (such as filtering then sorting), the index may change after each operation. Using it at the right time
reset_index()can keep the code predictable.
Notes
Other Extensions
Common Pandas Functions