Pandas pd.get_dummies() Function

Pandas 通用函数Common Pandas Functions


pd.get_dummies()It is a Pandas library function used forone-hot encoding of categorical variablesfunction. It converts categorical variables into binary (0/1) columns, with each category corresponding to a column.

One-hot encoding is a common technique in machine learning preprocessing because most algorithms cannot directly handle categorical data and need to convert it into numerical form.

Word Explanation: get_dummiesHere "dummy" means "dummy variable", referring to binary variables used in statistics and econometrics to represent categorical variables.


Basic Syntax and Parameters

pd.get_dummies()It is a top-level function in the Pandas library, used to convert categorical variables into one-hot encoding format.

Syntax Format

pd.get_dummies(data, prefix=None, prefix_sep='_', dummy_na=False, columns=None, drop_first=False, dtype=None)

Parameter Description

  • Parameter: data
    • Type: Series, DataFrame, or array-like object.
    • Description: The data to be one-hot encoded. Usually a Series or DataFrame containing categorical variables.
  • Parameter: prefix
    • Type: String, list of strings, or dictionary.
    • Description: The prefix for the generated new column names. If not specified, the original column names are used as the prefix.
  • Parameter: prefix_sep
    • Type: String.
    • Description: The separator between the prefix and the category name. Defaults to underscore'_'。
  • Parameter: dummy_na
    • Type: Boolean.
    • Description: IfTrue, it also creates a separate column for missing values (NaN). Defaults toFalse。
  • Parameter: columns
    • Type: List or None.
    • Description: The column names to encode. If not specified, all columns of type object, category, or boolean are encoded.
  • Parameter: drop_first
    • Type: Boolean.
    • Description: IfTrue, it will remove the first column of each categorical variable to avoid multicollinearity. In models such as logistic regression, it can avoid the dummy variable trap. Defaults toFalse。

Function Description

  • Return Value: Returns a DataFrame where each column corresponds to a category and the values are 0 or 1.
  • Effect: Converts categorical variables into numerical binary columns, making it easier for machine learning algorithms to process.

Examples

Let's, through a series of examples from simple to complex, thoroughly masterpd.get_dummies()the usage of pd.get_dummies().

Example 1: Basic Usage - One-Hot Encoding a Series

Example

import pandas as pd

# 1. Create a Series containing categorical variables
colors = pd.Series(['red', 'blue', 'green', 'red', 'green', 'blue'])

print("=== Original Series ===")
print(colors)

# 2. Use pd.get_dummies() for one-hot encoding
result = pd.get_dummies(colors)
print("\n"=== pd.get_dummies() one-hot encoding result ===")
print(result)

Expected output:

=== 原始 Series ===
0      red
1     blue
2     green
3      red
4     green
5     blue
dtype: object

=== pd.get_dummies() 独热编码结果 ===
    blue  green  red
0  False   False  True
1   True   False  False
2  False    True  False
3  False   False  True
4  False    True  False
5   True   False  False

Code explanation:

  1. The original Series contains three colors: red, blue, green.
  2. After one-hot encoding, each color becomes a column, usingTrue/Falseto indicate whether the row belongs to that category.
  3. Each row has exactly oneTrue, corresponding to the original color value.

Example 2: Encoding Specific Columns of a DataFrame

In data analysis, it is often only necessary to encode specific categorical columns while leaving numerical columns unchanged.

Example

import pandas as pd

# 1. Create a DataFrame containing numerical and categorical variables
df = pd.DataFrame({
    'name': ['Alice', 'Bob', 'Charlie', 'Diana'],
    'age': [25, 30, 35, 28],
    'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou'],
    'department': ['Sales', 'Engineering', 'Sales', 'HR']
})

print("=== Original DataFrame ===")
print(df)

# 2. Perform one-hot encoding on the specified categorical columns
result = pd.get_dummies(df, columns=['city', 'department'])
print("\n"=== One-hot encoding for city and department columns ===")
print(result)

Expected output:

=== 原始 DataFrame ===
      name  age       city  department
0    Alice   25    Beijing       Sales
1      Bob   30  Shanghai  Engineering
2  Charlie   35    Beijing       Sales
3   Diana   28  Guangzhou         HR

=== 对 city 和 department 列进行独热编码 ===
      name  age  city_Beijing  city_Guangzhou  city_Shanghai  department_Engineering  department_HR  department_Sales
0    Alice   25           True            False           False                False            True
1      Bob   30          False           False            True                True           False
2  Charlie   35           True           False           False                False            True
3   Diana   28          False            True           False                False           False

Code explanation:

  • Use thecolumnsparameter to specify that only thecityanddepartmentcolumns are encoded.
  • Numerical columnsageand text columnsnameremain unchanged.
  • The newly generated column names use the default separator underscore, such ascity_Beijing。

Example 3: Custom Prefix and Separator

You can use theprefixandprefix_sepparameter to customize the names of new columns.

<h2 class="example">Example
import pandas as pd

# 1. Create a DataFrame
df = pd.DataFrame({
    'color': ['red', 'blue', 'green', 'red'],
    'size': ['S', 'M', 'L', 'XL']
})

print("=== Original DataFrame ===")
print(df)

# 2. Use the prefix parameter to customize the prefix
result_prefix = pd.get_dummies(df, prefix=['color', 'size'])
print("\n"=== Using custom prefix ===")
print(result_prefix)

# 3. Use prefix_sep to customize the separator
result_sep = pd.get_dummies(df, prefix=['color', 'size'], prefix_sep='-')
print("\n"=== Using custom separator '-' ===")
print(result_sep)

Expected output:

=== 原始 DataFrame ===
   color  size
0    red     S
1   blue     M
2  green     L
3    red    XL

=== 使用自定义前缀 ===
   color_blue  color_green  color_red  size_L  size_M  size_S  size_XL
0       False        False       True   False   False    True   False
1        True       False      False   False    True   False   False
2       False        True      False    True   False   False   False
3       False        False       True   False   False   False    True

=== 使用自定义分隔符 '-' ===
   color-blue  color-green  color-red  size-L  ...

Code explanation:

  • prefix=['color', 'size']Specify different prefixes for different columns.
  • prefix_sep='-'Change the default underscore to a hyphen, and the new column names becomecolor-redthis form.

Example 4: Handling Missing Values and the drop_first Parameter

Example

import pandas as pd
import numpy as np

# 1. Data containing missing values
df = pd.DataFrame({
    'color': ['red', 'blue', np.nan, 'red', 'green'],
    'size': ['S', 'M', 'L', np.nan, 'XL']
})

print("=== DataFrame containing missing values ===")
print(df)

# 2. By default, missing values are not handled
result_default = pd.get_dummies(df, columns=['color'])
print("\n"=== Default: missing values not handled ===")
print(result_default)

# 3. dummy_na=True creates a separate column for missing values
result_na = pd.get_dummies(df, columns=['color'], dummy_na=True)
print("\n"=== dummy_na=True creates a column for missing values ===")
print(result_na)

# 4. drop_first=True removes the first column to avoid multicollinearity
result_drop = pd.get_dummies(df, columns=['size'], drop_first=True)
print("\n"=== drop_first=True removes the first column ===")
print(result_drop)

Expected output:

=== 包含缺失值的 DataFrame ===
   color  size
0    red     S
1   blue     M
2    NaN     L
3    red   NaN
4  green   XL

=== 默认不处理缺失值 ===
   size  color_blue  color_green  color_red
0     S       False        False       True
1     M        True       False      False
2     L       False        False      False
3     S       False        False       True
4    XL       False        True      False

=== dummy_na=True 为缺失值创建列 ===
   size  color_True  color_blue  color_green  color_red
0     S       False        False        False       True
1     M       False         True        False      False
2     L        True        False        False       False
3     S       False        False        False       True
4    XL       False        False         True       False

=== drop_first=True 删除第一列 ===
   color_blue  color_green  color_red
0       False        False       True
1        True       False      False
2       False        False      False
3       False        False       True
4       False        True       False


Code explanation:

  • dummy_na=TrueCreate an extra column for missing values (shown as the color_True column in this example).
  • drop_first=TrueRemoving the first column of each group (e.g., size_M is removed) can avoid multicollinearity issues in linear models.

Tip:One-hot encoding increases the data dimensionality (one column per category). If the number of categories is very large, it may lead to the curse of dimensionality. In such cases, you can consider using label encoding (pd.factorize()) or other encoding methods.

Pandas 常用函数Common Pandas Functions