)}"
pd.factorize()Is a Pandas library function used forencoding categorical variablesThe function converts categorical data into integer codes and returns a list of unique values.
It is a common technique in machine learning preprocessing, converting text categories into numeric values, making it easier for algorithms to process. Compared with one-hot encoding (pd.get_dummies()), factorization encoding does not increase the data dimension.
Word Definition: factorizeMeans "factorization", here referring to converting categorical variables into numeric factors (integer codes).
Basic Syntax and Parameters
pd.factorize()It is a top-level function of the Pandas library, used to encode categorical variables as integers.
Syntax Format
pd.factorize(values, sort=False, na_sentinel=-1, size_hint=None)
Parameter Description
- Parameter:
values- Type: Array-like object, such as Series, list, array, etc.
- Description: The categorical data to be encoded.
- Parameter:
sort- Type: Boolean value.
- If it is
True, sort the unique values before encoding. Default isFalse(encode in order of appearance).
na_sentinel
- Type: Integer.
- Integer used to mark missing values. Default is
-1. If set toNone, missing values will be kept as NaN.
Function Description
- Return Value: Returns a tuple (codes, uniques), where codes is the encoded integer array, and uniques is the array of unique values.
- Effect: Maps categorical variables to an integer sequence starting from 0, with each unique value corresponding to an integer.
Examples
Let's thoroughly master through a series of examples from simple to complexpd.factorize()its usage.
Example 1: Basic Usage - Encoding Categorical Variables
Example
import numpy as np
# 1. Create categorical data
colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow'])
print("=== Original categorical data ===")
print(colors)
# 2. Use pd.factorize() for encoding
codes, uniques = pd.factorize(colors)
print("\n=== pd.factorize() encoding result ===")
print(fEncoding: {codes}) # Encoded integer
print(fUnique values: {uniques}) # uniqueValueList
print(f"Type: {type(codes)}")
Expected output:
=== 原始分类数据 === 0 red 1 blue 2 green 3 red 4 blue 5 yellow 原始数据中 'red' 出现 2 次,'blue' 出现 2 次,'green' 和 'yellow' 各出现 1 次。 按出现顺序分配编码: - 'red' -> 0 - 'blue' -> 1 - 'green' -> 2 - 'yellow' -> 3 === 编码: [0 1 2 0 1 3] === 唯一值数组: ['red' 'blue' 'green' 'yellow']
Code explanation:
- The returned codes is an integer array, where each original value is replaced with the corresponding integer code.
- The uniques array stores all unique values, arranged in the corresponding order of encoding.
- By default, codes are assigned in order of appearance (the first appearing value is encoded as 0).
Example 2: sort Parameter - Sorting Unique Values Before Encoding
Usingsort=Trueyou can sort the unique values before encoding, making the results more predictable.
Example
import numpy as np
# Create categorical data
colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow'])
print(=== Raw Data ===)
print(colors)
# 1. No sorting (default) - encode in order of appearance
codes_unsorted, uniques_unsorted = pd.factorize(colors, sort=False)
print("\n=== sort=False (default) ===")
print(f"Codes: {codes_unsorted}")
print(fUnique values: {uniques_unsorted})
# 2. Encode after sorting - in alphabetical order
codes_sorted, uniques_sorted = pd.factorize(colors, sort=True)
print("\n=== sort=True (in alphabetical order) ===)
print(fCodes: {codes_sorted})
print(f"Unique values: {uniques_sorted}")
# 3. Sort numeric values
numbers = pd.Series([5, 2, 8, 5, 3, 2, 9])
codes_num, uniques_num = pd.factorize(numbers, sort=True)
print("\n=== Numeric Sorting Encoding ===)
print(f"Codes: {codes_num}")
print(fUnique values: {uniques_num})
Expected output:
=== 原始数据 === 0 red 1 blue 2 green 3 red 4 blue 5 yellow === sort=False (默认) === 编码: [0 1 2 0 1 3] 唯一值: ['red' 'blue' 'green' 'yellow'] === sort=True (按字母顺序) === 编码: [2 0 1 2 0 3] 唯一值: ['blue' 'green', 'red', 'yellow'] 按字母顺序排列:blue, green, red, yellow - blue -> 0 - green -> 1 - red -> 2 - yellow -> 3 === 数值排序编码 === 编码: [2 0 3 2 1 0 4] 唯一值: [2, 3, 5, 8, 9] 按数值排序:2, 3, 5, 8, 9 - 2 -> 0 - 3 -> 1 - 5 -> 2 - 8 -> 3 - 9 -> 4
Code explanation:
sort=TrueAfter sorting the unique values and then encoding, the coding order is deterministic and predictable.- This is useful when reproducible results are required.
Example 3: Handling Missing Values
na_sentinelThe parameter is used to control the handling of missing values.
Example
import numpy as np
Data containing missing values
data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', 'b'])
print("=== Data with missing values ===")
print(data)
# 2. Default behavior (missing value → -1)
codes_default, uniques_default = pd.factorize(data)
print("\n=== Default handling (missing values = -1) ===)
print(fEncoding: {codes_default})
print(f"Unique values: {uniques_default}")
# 3. Custom missing value encoding (simulating na_sentinel=-999)
codes_custom = np.where(codes_default == -1, -999, codes_default)
print("\n=== Custom missing value encoding (-999) ===)
print(fCodes: {codes_custom})
# 4. Keep NaN (simulate na_sentinel=None)
codes_nan = codes_default.astype('float')
codes_nan[codes_nan == -1] = np.nan
print("\n=== Keep NaN ===)
print(f"Codes: {codes_nan}")
print(f"Encoding type: {type(codes_nan)
Expected output:
=== 包含缺失值的数据 === 0 a 1 b 2 NaN 3 a 4 None 5 c 6 b dtype: object === 默认处理(缺失值 = -1)=== 编码: [ 0 1 -1 0 -1 2 1] 唯一值: Index(['a', 'b', 'c'], dtype='object') === 自定义缺失值编码(-999)=== 编码: [ 0 1 -999 0 -999 2 1] === 保持 NaN === 编码: [ 0. 1. nan 0. nan 2. 1.] 编码类型: <class 'numpy.float64'>
Example 4: Application in Machine Learning
Factorization encoding is very practical in machine learning preprocessing, especially for high-cardinality categorical variables.
Example
Create a DataFrame containing categorical variables
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Charlie', 'Diana', 'Eve', 'Frank'],
'city': ['Beijing', 'Shanghai', 'Beijing', 'Guangzhou', 'Shanghai', 'Beijing'],
'department': ['Sales', 'Engineering', 'Sales', 'HR', 'Engineering', 'Sales']
})
print("=== Original DataFrame ===")
print(df)
# 2. Factorize categorical columns
# Create a mapping dictionary
city_codes, city_uniques = pd.factorize(df['city'])
dept_codes, dept_uniques = pd.factorize(df['department'])
print("\n=== Factorization Encoding Result ===)
print(f"city codes: {city_codes.tolist()}")
print(f"city unique values: {city_uniques}")
print(f"department code: {dept_codes.tolist()}")
print(f"department unique values: {dept_uniques}")
# 3. Create the encoded DataFrame
df_encoded = df.copy()
df_encoded['city_encoded'] = city_codes
df_encoded['dept_encoded'] = dept_codes
print("\n=== Encoded DataFrame ===)
print(df_encoded)
# 4. Use a mapping dictionary to encode new data
print("\n=== Mapping dictionary ===)
city_mapping = dict(zip(city_uniques, range(len(city_uniques))))
print(fcity mapping: {city_mapping})
# Encode new data
new_cities = pd.Series(['Beijing', 'Shanghai', 'Shenzhen'])
new_codes = new_cities.map(city_mapping)
print(f"\nNew data encoding: {new_codes.tolist()}")
Expected output:
=== 原始 DataFrame ===
name city department
0 Alice Beijing Sales
1 Bob Shanghai Engineering
2 Charlie Beijing Sales
3 Diana Guangzhou HR
4 eve Shanghai Engineering
5 Frank Beijing Sales
=== 因子化编码结果 ===
city 编码: [0 1 0 2 1 0]
city 唯一值: ['Beijing' 'Shanghai' 'Guangzhou']
department 编码: [0 1 0 2 1 0]
department 唯一值: ['Sales' 'Engineering' 'HR']
=== 编码后的 DataFrame ===
name city department city_encoded dept_encoded
0 Alice Beijing Sales 0 0
1 Bob Shanghai Engineering 1 1
2 Charlie Beijing Sales 0 0
3 Diana Guangzhou HR 2 2
4 Eve Shanghai Engineering 1 1
5 Frank Beijing Sales 0 0
=== 映射字典 ===
city 映射: {'Beijing': 0, 'Shanghai': 1, 'Guangzhou': 2}
Code explanation:
pd.factorize()Returns a tuple that can be directly unpacked to obtain the codes and unique values.- After creating a mapping dictionary, the same encoding can be applied to new data, facilitating machine learning model inference.
- Factorization encoding does not increase the data dimension, making it suitable for handling high-cardinality categorical variables.
Other ExtensionsTip:
pd.factorize()Suitable for handling ordered or unordered categorical variables. If one-hot encoding is needed (for algorithms such as logistic regression), please usepd.get_dummies()function.
Pandas Common Functions