Pandas df.rename() function

Pandas 常用函数Common Pandas functions


df.rename()is a function in Pandas used to rename the column names or index of a DataFrame.

In data analysis, clear column names and indexes are very important for code readability and maintainability.rename()It allows you to flexibly modify column names and index names, making data more standardized and easier to understand. This is especially useful during data import, cleaning, and report generation.


Basic syntax and parameters

rename()is a member function of DataFrame, called via the dot operator.to invoke.

Syntax format

DataFrame.rename(mapper=None, index=None, columns=None, axis=None, inplace=False, errors='ignore')

Parameter description

Parameter Type Required Description Default value
mapper dict or function Optional Mapping rules for renaming an axis (row index or column names). Can be a dictionary or a function. Used with theaxisparameter. None
index dict or function Optional Directly specify renaming rules for the row index. Takes precedence overmapperandaxis。 None
columns dict or function Optional Directly specify renaming rules for column names. Takes precedence overmapperandaxis。 None
axis int or str Optional Specifies the axis to rename.0or'index'represents the row index;1or'columns'represents column names. None
inplace bool Optional If set toTrue, it modifies the original DataFrame directly and returns no new object; if set toFalse, a new DataFrame is returned, and the original data remains unchanged. False
errors str Optional Controls error handling.'raise'means an exception is raised when the corresponding key is not found;'ignore'means ignoring non-existent keys. 'ignore'

Return value description

  • Returns a new DataFrame (ifinplace=False), orNone(ifinplace=True)。
  • In the returned DataFrame, the column names or index have been renamed.

Examples

Let us thoroughly master through a series of examplesrename()the usage of.

Example 1: Rename column names

Use thecolumnsparameter to conveniently rename column names.

Example

import pandas as pd

# Create a DataFrame
data = {
    'name': ['Zhang San', 'Li Si', 'Wang Wu'],
    'age': [25, 30, 35],
    'salary': [5000, 6000, 7000]
}
df = pd.DataFrame(data)

print("Original data:")
print(df)
print("=" * 50)

# Use a dictionary to rename column names
df_renamed = df.rename(columns={
    'name': 'Name',
    'age': 'age',
    'salary': 'Salary'
})

print("Data after renaming column names:")
print(df_renamed)

Expected output:

原始数据:
   name  age  salary
0  张三   25   5000
1  李四   30   6000
2  王五   35   7000
==================================================
重命名列名后的数据:
    姓名  年龄   薪资
0  张三   25  5000
1  李四   30  6000
2  王五   35  7000

Code explanation:

  1. The original data's column names are in English, and we use a dictionary to rename them to Chinese.
  2. Using thecolumnsparameter only affects column names, not the row index.
  3. This is a common data normalization operation.

Example 2: Rename row index

Use theindexparameter to rename the row index.

Example

import pandas as pd

# Create a DataFrame with default index
data = {
    Fruit: [Apple, Banana, Orange],
    'Quantity': [10, 20, 15]
}
df = pd.DataFrame(data)

print("Original data (default index):")
print(df)
print("=" * 50)

# Use a dictionary to rename row index
df_renamed = df.rename(index={
    0: 'First line',
    1: 'Second line',
    2: 'Third line'
})

print("Data after renaming row index:")
print(df_renamed)

Expected output:

原始数据(默认索引):
   水果  数量
0  苹果  10
1  香蕉  20
2  橙子  15
==================================================
重命名行索引后的数据:
     水果  数量
第一行   苹果  10
第二行   香蕉  20
第三行   橙子  15

Code explanation:

  • The original data's row index is numbers like 0, 1, 2.
  • Using theindexparameter can rename them to more meaningful names.
  • This is especially useful when generating reports.
  • Example 3: Rename both column names and row index

    It can rename column names and row index at the same time.

    Example

    import pandas as pd

    # Create a DataFrame
    data = {
        'A': [1, 2, 3],
        'B': [4, 5, 6]
    }
    df = pd.DataFrame(data, index=['x', 'y', 'z'])

    print("Original data:")
    print(df)
    print("=" * 50)

    # Rename both column names and row index
    df_renamed = df.rename(columns={'A': Column A, 'B': Column B}, index={'x': Row X, 'y': Row Y, 'z': Row Z})

    print("Data after renaming:")
    print(df_renamed)

    Expected output:

    原始数据:
       A  B
    x  1  4
    y  2  5
    z  3  6
    ==================================================
    重命名后的数据:
        列A  列B
    行X   1   4
    行Y   2   5
    行Z   3   6
    

    Code explanation:

    • You can use bothcolumnsandindexparameters.
    • This is very convenient during data import and standardization.
    • Example 4: Using a function to rename

      rename()It also accepts a function as a parameter, allowing batch transformation of column names or index.

      Example

      import pandas as pd

      # Create a DataFrame with relatively complex column names
      data = {
          'USER_NAME': ['Zhang San', 'Li Si'],
          'USER_AGE': [25, 30],
          'USER_SALARY': [5000, 6000]
      }
      df = pd.DataFrame(data)

      print("Original data:")
      print(df)
      print("=" * 50)

      # Use a function to convert column names to lowercase
      df_lower = df.rename(columns=str.lower)

      print("After converting column names to lowercase:")
      print(df_lower)
      print("=" * 50)

      # Use a function to remove prefix (e.g., USER_)
      df_clean = df.rename(columns=lambda x: x.replace('USER_', ''))

      print("After removing prefix:")
      print(df_clean)
      print("=" * 50)

      # Use a function to add prefix
      df_prefix = df.rename(columns=lambda x: 'col_' + x)

      print("After adding prefix:")
      print(df_prefix)

      Expected output:

      原始数据:
        USER_NAME  USER_AGE  USER_SALARY
      0       张三       25       5000
      1       李四       30       6000
      列名转小写后:
        user_name  user_age  user_salary
      0       张三       25       数据清洗
      1       李四       30       6000
      ==================================================
      去除前缀后:
        NAME  AGE  SALARY
      0  张三   25  5000
      1  李四   30  6000
      

      Code explanation:

      • str.loweris the lowercase method for strings.
      • Usinglambdafunctions can flexibly handle complex naming rules.
      • This method is suitable for batch processing a large number of column names.
      • Example 5: Using inplace parameter for in-place modification

        Useinplace=Trueto modify the original DataFrame directly without returning a new object.

        Example

        import pandas as pd

        # Create a DataFrame
        data = {
            'name': ['Zhang San', 'Li Si'],
            'age': [25, 30]
        }
        df = pd.DataFrame(data)

        print("Original data:")
        print(df)
        print(f"Original data id: {id(df)}")
        print("=" * 50)

        # Use inplace=True to modify in place
        df.rename(columns={'name': 'Name', 'age': 'age'}, inplace=True)

        print("Data after modifying with inplace=True:")
        print(df)
        print(f"Modified data id: {id(df)}")

        Expected output:

        原始数据:
           name  age
        0  张三   25
        1  李四   30
        原始数据 id: 140234567890
        ==================================================
        使用 inplace=True 修改后的数据:
            姓名  年龄
        0  张三   25
        1  李四   30
        修改后数据 id: 140234567890  # 同一个对象
        

        Code explanation:

        • Usinginplace=Trueafter, the memory address (id) of the DataFrame remains unchanged, indicating an in-place modification.
        • This method can save memory, especially when working with large datasets.
        • Example 6: Handling non-existent column names

          Use theerrorsparameter to control behavior when column names do not exist.

          Example

          import pandas as pd

          # Create a DataFrame
          data = {
              'A': [1, 2, 3],
              'B': [4, 5, 6]
          }
          df = pd.DataFrame(data)

          print("Original data:")
          print(df)
          print("=" * 50)

          # Attempt to rename a non-existent column, using errors='ignore' (default)
          df_renamed = df.rename(columns={'C': 'C_new'})  # Column C does not exist

          print("Renaming non-existent column (C) with default behavior:")
          print(df_renamed)  # No change will occur
          print("=" * 50)

          # Attempt to rename a non-existent column, using errors='raise'
          try:
              df_renamed2 = df.rename(columns={'C': 'C_new'}, errors='raise')
          except KeyError as e:
              print(f"Exception raised: {e}")

          Expected output:

          原始数据:
             A  B
          0  1  4
          1  2  5
          2  3  6
          ==================================================
          重命名不存在的列(C),使用默认行为:
             A  B
          0  1  4
          1  2  5
          2  3  6  # 无变化
          ==================================================errors
          ==================================================
          抛出异常: 'C'
          

          Code explanation:

          • By default (errors='ignore'), if a column does not exist, no change will occur and no error will be raised.
          • If set toerrors='raise', an exception will be raised.KeyErrorexception.
          • Example 7: Using the axis parameter

            UseaxisThe parameter allows more flexible specification of the axes to rename.

            Example

            import pandas as pd

            # Create a DataFrame
            data = {
                'col1': [1, 2, 3],
                'col2': [4, 5, 6]
            }
            df = pd.DataFrame(data, index=['row1', 'row2', 'row3'])

            print("Original data:")
            print(df)
            print("=" * 50)

            # Use axis='columns' to rename columns
            df_renamed_cols = df.rename({'col1': 'First column', 'col2': 'Second column'}, axis='columns')

            print("After renaming columns with axis='columns':")
            print(df_renamed_cols)
            print("=" * 50)

            # Use axis='index' to rename row index
            df_renamed_idx = df.rename({'row1': 'First row', 'row2': 'Second row', 'row3': 'Third row'}, axis='index')

            print("After renaming rows with axis='index':")
            print(df_renamed_idx)

            Expected output:

            原始数据:
                  col1  col2
            row1     1     4
            row2     2     5
            row3     3     6
            ==================================================
            使用 axis='columns' 重命名列后:
                  第一列  第二列
            row1     1     4
            row2     2     5
            row3     3     6
            

            Code explanation:

            • axis='columns'oraxis=1Means operating on columns.
            • axis='index'oraxis=0Means operating on row indices.
            • This syntax is consistent with NumPy's style.

            • Notes

              • rename()By default, the original DataFrame is not modified. To modify it in place, useinplace=Trueparameter.
              • Prefer usingcolumnsandindexparameter instead ofmapperandaxis, because the former is clearer and easier to understand.
              • When using a function for renaming, make sure the function returns a string type, otherwise it may cause errors.
              • When working with large datasets, usinginplace=Truecan save memory.
              • Before renaming, it is recommended to check the current column names. You can usedf.columnsto view.

              Pandas 常用函数Common Pandas Functions

              Other Extensions