Pandas df.reset_index() Function

Pandas 常用函数Common Pandas Functions


df.reset_index()is a function in Pandas used to reset the row index of a DataFrame.

During data processing, the index may become discontinuous or fail to meet your needs.reset_index()It can reset the index to the default integer index (0, 1, 2, ...), or convert the current index into a column. This is very useful for data cleaning, sorting, and organizing data after filtering.


Basic Syntax and Parameters

reset_index()is a member function of DataFrame, invoked via the dot operator.to call.

Syntax Format

DataFrame.reset_index(level=None, drop=False, inplace=False, col_level=0, col_fill='')

Parameter Description

Parameter Type Required Description Default Value
level int, str, tuple, or list Optional Specifies the index level(s) to reset. For a MultiIndex, you can choose to reset some levels. By default, all levels are reset. None
drop bool Optional If it isTrue, discard the original index without converting it to a column; if it isFalse, keep the original index as a new column. False
inplace bool Optional If it isTrue, modify the original DataFrame directly and do not return a new object; if it isFalse, return a new DataFrame and keep the original data unchanged. False
col_level int or str Optional If the column names are a MultiIndex, specify which level to insert the index into. 0
col_fill str Optional If the column names are a MultiIndex, this is used to fill in the column names of other levels when the original index is converted to a column. ''

Return Value Description

  • Returns a new DataFrame (ifinplace=False), orNone(ifinplace=True)。
  • In the returned DataFrame, the index has been reset. The original index (ifdrop=False) will be kept as a new column.

Examples

Let's thoroughly master, through a series of examples,reset_index()the usage of.

Example 1: Basic Usage - Reset the Index

The most common usage is to reset the index to the default integer index.

Example

import pandas as pd
import numpy as np

# Create a DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'Score': [85, 92, 78]
}
df = pd.DataFrame(data)

print(Original data (default index):)
print(df)
print("=" * 50)

# Sort the data, resulting in a non-contiguous index
df_sorted = df.sort_values(by='Score', ascending=False)

print(Sorted data (non-contiguous indices):)
print(df_sorted)
print("=" * 50)

# Reset index
df_reset = df_sorted.reset_index()

print(After resetting index:)
print(df_reset)

Expected output:

原始数据(默认索引):
   姓名  成绩
0  张三   85
1  李四   92
2  王五   78
==================================================
排序后的数据(索引不连续):
   姓名  成绩
1  李四   92
0  张三   85
2  王五   78
==================================================
重置索引后:
   姓名  成绩  index
0  李四   92       1  # 原来的索引作为新列
1  张三   85       0
2  王五   78       2

Code explanation:

  1. After sorting, the index becomes 1, 0, 2, no longer continuous 0, 1, 2.
  2. After usingreset_index(), the index is reset to 0, 1, 2.
  3. The original index is kept as a new column, "index".

Example 2: Use the drop Parameter to Discard the Original Index

If you do not need to keep the original index, you can usedrop=Trueparameter.

Example

import pandas as pd
import numpy as np

# Create a DataFrame
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'Score': [85, 92, 78]
}
df = pd.DataFrame(data)

# Sort the Data
df_sorted = df.sort_values(by='Score', ascending=False)

print("Sorted Data:")
print(df_sorted)
print("=" * 50)

# Reset index, drop the original index
df_reset = df_sorted.reset_index(drop=True)

print(After resetting index (dropping original index):)
print(df_reset)

Expected output:

排序后的数据:
   姓名  成绩
1  李四   92
0  张三   85
2  王五   78
==================================================
重置索引(丢弃原索引)后:
   姓名  成绩
0  李四   92
1  张三   85
2  王五   78

Code explanation:

  • After usingdrop=True, the original index is directly discarded and will not be kept as a new column.
  • This is useful when you do not need to keep the original index information.

Example 3: Reset the Index After Filtering Data

After data filtering, the index may be discontinuous. Usereset_index()to tidy up the data.

Example

import pandas as pd
import numpy as np

# Create a DataFrame containing all student grades
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', 'Qian Qi'],
    'Score': [85, 92, 78, 90, 88]
}
df = pd.DataFrame(data)

print(“Raw data:”)
print(df)
print("=" * 50)

# Filter students with scores greater than 85
df_filtered = df[df['Score'] > 85]

print(Filtered data (non-contiguous indices):)
print(df_filtered)
print("=" * 50)

# Reset index
df_reset = df_filtered.reset_index(drop=True)

print(After resetting index:)
print(df_reset)

Expected output:

原始数据:
   姓名  成绩
0  张三   85
1  李四   92
2  王五   78
3  赵六   90
4  钱七   88
==================================================
筛选后的数据(索引不连续):
   姓名  成绩
1  李四   92
3  赵六   90
4  钱七   88

Code explanation:

  • After filtering, the retained indices are the original 1, 3, 4, not the continuous 0, 1, 2.
  • Usingreset_index(drop=True)can give a continuous index, which is more convenient for subsequent processing.
  • Example 4: Reset the Index After Removing Duplicate Rows

    After deleting duplicate data, you can also usereset_index()to tidy up the index.

    Example

    import pandas as pd
    import numpy as np

    # Create a DataFrame with duplicate rows
    data = {
        'Name': ['Zhang San', 'Li Si', 'Zhang San', 'Wang Wu'],
        'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou']
    }
    df = pd.DataFrame(data)

    print(Original data (including duplicate rows):)
    print(df)
    print("=" * 50)

    # Deleterepeatline
    df_unique = df.drop_duplicates()

    print(After removing duplicate rows (non-contiguous indices):)
    print(df_unique)
    print("=" * 50)

    # Reset index
    df_reset = df_unique.reset_index(drop=True)

    print(After resetting index:)
    print(df_reset)

    Expected output:

    原始数据(包含重复行):
       姓名  城市
    0  张三  北京
    1  李四  上海
    2  张三  北京  # 重复
    3  王五  广州
    ==================================================
    删除重复行后(索引不连续):
       姓名  城市
    0  张三  北京
    1  李四  上海
    3  王五  广州
    

    Code explanation:

    • After deleting duplicate rows, the retained indices are 0, 1, 3, not continuous 0, 1, 2.
    • After resetting the index, the data is tidier.
    • Example 5: Reset the Index After Handling Missing Values

      After deleting rows with missing values, the index can also be reset.

      Example

      import pandas as pd
      import numpy as np

      # Create a DataFrame with missing values
      data = {
          'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
          'Score': [85, np.nan, 92, 78]
      }
      df = pd.DataFrame(data)

      print(“Raw data:”)
      print(df)
      print("=" * 50)

      # Delete rows with missing values
      df_cleaned = df.dropna()

      print(After removing missing values (non-contiguous indices):)
      print(df_cleaned)
      print("=" * 50)

      # Reset index
      df_reset = df_cleaned.reset_index(drop=True)

      print(After resetting index:)
      print(df_reset)

      Expected output:

      原始数据:
         姓名   成绩
      0  张三  85.0
      1  李四   NaN  # 缺失值
      2  王五  92.0
      3  赵六  78.0
      ==================================================
      删除缺失值后(索引不连续):
         姓名   成绩
      0  张三  85.0
      2  王五  92.0
      3  赵六  78.0
      

      Code explanation:

      • After deleting row 1 (the score of Li Si is NaN), the retained indices are 0, 2, 3.
      • Usingreset_index()can give a continuous index of 0, 1, 2.
      • Example 6: Use the inplace Parameter to Modify In Place

        Usinginplace=Trueyou can directly modify the original DataFrame.

        Example

        import pandas as pd
        import numpy as np

        # Create a DataFrame
        data = {
            'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
            'Score': [85, 92, 78]
        }
        df = pd.DataFrame(data)

        # Filter data
        df_filtered = df[df['Score'] > 80]

        print(Filtered data (index not reset):)
        print(df_filtered)
        print(fData id: {id(df_filtered)})
        print("=" * 50)

        # Use inplace=True to reset the index in place
        df_filtered.reset_index(drop=True, inplace=True)

        print(After resetting with inplace=True:)
        print(df_filtered)
        print(fData id: {id(df_filtered)})

        Expected output:

        筛选后的数据(未重置):
           姓名  成绩
        1  李四   92
        0  张三   85
        数据 id: 140234567890
        ==================================================
        使用 inplace=True 重置后:
           姓名  成绩
        0  李四   92
        1  张三   85
        数据 id: 140234567890  # 同一个对象
        

        Code explanation:

        • After usinginplace=True, the memory address (id) of the DataFrame remains unchanged.
        • This method can save memory.
        • Example 7: Resetting a Multi-level Index

          For a MultiIndex, you can reset specific levels or all levels.

          Example

          import pandas as pd
          import numpy as np

          # Create a DataFrame with a MultiIndex
          df = pd.DataFrame({
              'Score': [85, 92, 78, 88, 90, 95]
          }, index=pd.MultiIndex.from_tuples([
              ('Class 1', 'Zhang San'), ('Class 1', 'Li Si'), ('Class 1', 'Wang Wu'),
              ('Class 2', 'Zhao Liu'), ('Class 2', 'Qian Qi'), ('Class 2', 'Sun Ba')
          ], names=['Class', 'Name']))

          print(Original data (MultiIndex):)
          print(df)
          print("=" * 50)

          # Reset all index levels
          df_reset_all = df.reset_index()

          print(Reset all index levels:)
          print(df_reset_all)
          print("=" * 50)

          # Only reset the first index level (Class)
          df_reset_first = df.reset_index(level=0)

          print(Only reset the Class index:)
          print(df_reset_first)

          Expected output:

          原始数据(多层索引):
                     成绩
          班级  姓名
          一班  张三     85
               李四     92
               王五     78
          二班  赵六     88
               钱七     90
               孙八     95
          ==================================================
          重置所有索引级别:
             班级  姓名  成绩
          0  一班  张三   85
          1  一班  李四   92
          2  一班  王五   78
          3  二班  赵六   88
          4  二班  钱七   90
          5  二班  钱七   95
          

          Code explanation:

          • A DataFrame with a MultiIndex can usereset_index()to convert the index to columns.
          • Use thelevelparameter to choose which level of the index to reset.

          • Notes

            • reset_index()The original DataFrame is not modified by default. To modify it in place, use theinplace=Trueparameter.
            • By default (drop=False), the original index is kept as a new column. If you do not need to keep it, usedrop=True。
            • Before usingreset_index(), ensure that there is no important index information to retain.
            • For time series data, you may need to reset the time index after resetting the index.
            • In chained operations (such as filtering then sorting), the index may change after each operation. Using it at the right timereset_index()can keep the code predictable.

            Pandas 常用函数Pandas Common Functions

            Other Extensions