Pandas pd.unique() Function
pd.unique()is used in the Pandas library toget unique values from an arrayfunction. It returns all non-duplicate values from the input array, removing duplicates.
This is a common operation in data analysis, such as counting how many distinct values are in a column, or getting all categories of a categorical variable.
Word Definition: uniqueIt means "unique, distinctive"; here it refers to returning non-duplicate values from the array.
Basic Syntax and Parameters
pd.unique()is a top-level function of the Pandas library, used to extract unique values from an array.
Syntax Format
pd.unique(values)
Parameter Description
- Parameter:
values- Type: array-like object, such as a list, Series, one-dimensional array, etc.
- Description: The input data from which to extract unique values. It can be any one-dimensional array structure.
Function Description
- Return Value: Returns an ndarray (NumPy array) containing all unique values.
- Effect: Removes duplicate values from the input data, keeping each value only once.
Examples
Let's use a series of examples from simple to complex to thoroughly masterpd.unique()its usage.
Example 1: Basic Usage - Extracting Unique Values from a Series
Example
import numpy as np
# 1. Create a Series containing duplicate values
colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow', 'red'])
print(=== Original Series ===)
print(colors)
# 2. Use pd.unique() to get unique values
unique_values = pd.unique(colors)
print("\n=== pd.unique() Unique Values ===")
print(unique_values)
print(f"\nNumber of unique values: {len(unique_values)}")
Expected output:
=== 原始 Series === 0 red 1 blue 2 green 3 red 4 blue 5 yellow 6 red dtype: object === pd.unique() 唯一值 === ['red' 'blue' 'green' 'yellow'] 唯一值数量: 4
Code Explanation:
- The original Series has 7 elements, but only 4 unique values.
pd.unique()Returns a NumPy array containing all non-duplicate values.- The returned result does not guarantee order (but it usually follows the order of first appearance).
Example 2: Extracting Unique Values from Lists and Arrays
pd.unique()It can not only handle Series, but also various array structures.
Example
import numpy as np
# 1. Extract unique values from a list
numbers = [1, 2, 3, 2, 1, 4, 5, 3, 2]
print(=== Original List ===)
print(numbers)
unique_numbers = pd.unique(numbers)
print("\n=== Unique Values from the List ===")
print(unique_numbers)
print(f"Type: {type(unique_numbers)}")
# 2. Extract unique values from a NumPy array
arr = np.array(['a', 'b', 'a', 'c', 'b', 'd'])
print("\n=== NumPy Array ===")
print(arr)
print("Unique values:", pd.unique(arr))
# 3. Extract from a DataFrame column
df = pd.DataFrame({
'name': ['Alice', 'Bob', 'Charlie', 'Alice', 'Diana'],
'city': ['Beijing', 'Shanghai', 'Beijing', 'Beijing', 'Guangzhou']
})
print("\n=== DataFrame ===")
print(df)
print("\nUnique values of the name column:, pd.unique(df['name']))
print("Unique values of the city column:", pd.unique(df['city']))
Expected output:
=== 原始列表 ===
[1, 2, 3, 2, 1, 4, 5, 3, 2]
=== 列表的唯一值 ===
[1 2 3 4 5]
Type: <class 'numpy.ndarray'>
=== NumPy 数组 ===
['a' 'b' 'a' 'c' 'b' 'd']
唯一值: ['a' 'b' 'c' 'd']
=== DataFrame ===
name city
0 Alice Beijing
1 Bob Shanghai
2 Charlie Beijing
3 Alice Beijing
4 Diana Guangzhong
=== name 列的唯一值: ['Alice' 'Bob' 'Charlie' 'Diana']
=== city 列的唯一值: ['Beijing' 'Shanghai' 'Guangzhou']
Code Explanation:
pd.unique()It can handle Python lists, NumPy arrays, and DataFrame columns.- The returned result is always a one-dimensional NumPy array.
- When working with a DataFrame, you need to specify the specific column (using bracket syntax).
Example 3: Handling Numerical Data and Sorting
When dealing with numerical unique values, you can conveniently perform sorting and statistical analysis.
Example
import numpy as np
# 1. A Series containing duplicate numerical values
scores = pd.Series([85, 90, 78, 85, 92, 90, 78, 88, 95])
print(=== Original Scores ===)
print(scores)
# 2. Get unique values and sort them
unique_scores = pd.unique(scores)
print("\n=== Unique Values (Unsorted) ===)
print(unique_scores)
# 3. Sorted unique values
unique_sorted = np.unique(unique_scores)
print("\n=== Unique Values (Sorted) ===)
print(unique_sorted)
# 4. Count the occurrences of each unique value
print("\n=== Unique Values and Frequencies ===)
value_counts = pd.Series(scores).value_counts()
print(value_counts)
# 5. Get statistical information for the unique values
print("\n=== Basic Statistics ===)
print(f"Number of unique values: {len(unique_scores)}")
print(f"Minimum value: {unique_scores.min()}")
print(f"Maximum value: {unique_scores.max()}")
print(f"Average value: {unique_scores.mean():.2f}")
Expected output:
>--- 原始分数 --- 0 85 1 90 2 78 85、90、78 各出现 2 次,88、95 各出现 1 次。 === 唯一值及频次 === 85 2 90 2 78 2 88 1 95 1 dtype: int === 基本统计 === 唯一值数量: 5 最小值: 78 最大值: 95 平均值: 86.40
Code Explanation:
- You can use
np.unique()to sort the results. - Combined with
value_counts()you can see the frequency of each unique value. - The unique value array can be used for numerical calculations like a normal array.
Example 4: Handling Missing Values
pd.unique()It will return missing values (NaN) as a valid unique value.
Example
import numpy as np
# 1. Data containing missing values
data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', np.nan, 'b'])
print(=== Data with Missing Values ===)
print(data)
# 2. Get unique values (NaN will be treated as a unique value)
unique_with_nan = pd.unique(data)
print("\n=== Unique Values Including NaN ===)
print(unique_with_nan)
print(f"Number of unique values (including NaN): {len(unique_with_nan)}")
# 3. Unique values after excluding NaN
unique_no_nan = pd.unique(data[data.notna()])
print("\n=== Unique Values Excluding NaN ===)
print(unique_no_nan)
print(f"Number of unique values (excluding NaN): {len(unique_no_nan)}")
# 4. Use isna() to check whether NaN exists
has_nan = pd.isna(unique_with_nan).any()
print(f"\nWhether NaN exists: {has_nan}")
Expected output:
=== 包含缺失值的数据 ===
0 a
1 b
2 NaN
3 a
4 None
5 different 'c' appears twice: None and NaN are both treated as missing values.
print("\n=== 唯一值数量(含 NaN): 3 ===")
print("['a', 'b', nan]")
=== 排除 NaN 后的唯一值 ===
['a', 'b', 'c']
唯一值数量(不含 NaN): 3
Code Explanation:
- Both NaN and None are treated as missing values and count as only one in the unique values.
- You can use
data[data.notna()]to first filter out missing values, then extract unique values. pd.isna()You can check whether the unique value array contains NaN.
Other ExtensionsTip:
pd.unique()The return value is a NumPy array. If you need to return a pandas object (such as a Series), you can useSeries.unique()method, which will return a Series.
Common Pandas Functions