Pandas pd.value_counts() Function
pd.value_counts()is a function in the Pandas library used forcounting the frequency of each value. It calculates the number of occurrences of each unique value in the array and returns them sorted by frequency in descending order.
This is one of the most commonly used functions in categorical data analysis, allowing you to quickly understand the distribution of the data, such as counting voting results, product category distributions, etc.
Word Meaning: value_countsIt means "count values", i.e., counting the number of occurrences of each value.
Basic Syntax and Parameters
pd.value_counts()It is a top-level function of the Pandas library, used to calculate the frequency of each unique value.
Syntax Format
pd.value_counts(values, sort=True, ascending=False, normalize=False, bins=None, dropna=True)
Parameter Description
- Parameter:
values- Type: Series, array-like object.
- Description: The data whose frequencies are to be counted. Usually a Series.
- Parameter:
sort- Type: Boolean.
- Whether to sort by frequency. Defaults to
True(descending order by frequency).
ascending
- Type: Boolean.
- If
True, sort by frequency in ascending order; ifFalse(default), sort by frequency in descending order.
normalize
- Type: Boolean.
- If
True, return the proportion of each value (between 0 and 1) instead of the absolute frequency. Defaults toFalse。
bins
- Type: Integer or None.
- If an integer is specified, the numeric data is binned and frequencies are counted (similar to a histogram). Not applicable to categorical data.
dropna
- Type: Boolean.
- Whether to include NaN in the statistics results. Defaults to
True(not included).
Function Description
- Return Value: Returns a Series where the index is the unique values and the values are the frequencies (or proportions).
- Effect: Counts and displays the number of occurrences of each unique value.
Examples
Let's thoroughly masterpd.value_counts()the usage of.
Example 1: Basic Usage - Counting Frequencies of Categorical Data
Example
import numpy as np
# 1. Create a Series containing duplicate values
colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow', 'red', 'blue', 'blue'])
print("=== Original Series ===")
print(colors)
# 2. Use pd.value_counts() to count frequencies
result = pd.value_counts(colors)
print("\n=== pd.value_counts() frequency statistics ===")
print(result)
Expected output:
=== 原始 Series === 0 red 1 blue 2 green 3 red 4 blue 5 yellow 6 red 7 blue 8 blue dtype: object === pd.value_counts() 频次统计 === blue 4 red 3 green 1 yellow 1 dtype: int64
Code explanation:
- The results are sorted by frequency in descending order; blue appears the most (4 times), red comes next (3 times), and green and yellow each appear 1 time.
- What is returned is a Series, with the color values as the index and the counts as the values.
Example 2: Using the normalize Parameter to Calculate Proportions
normalize=TrueFrequencies can be converted into proportions, making it easier to view the relative distribution of the data.
Example
import numpy as np
# Create voting data
votes = pd.Series(['A', 'B', 'A', 'C', 'A', 'B', 'A', 'A', 'C', 'B', 'A', 'B'])
print("=== Voting data ===")
print(votes)
# 1. Absolute frequency
print("\n=== Absolute frequency ===")
print(pd.value_counts(votes))
# 2. Relative proportions (normalize=True)
print("\n=== Relative proportions ===")
result_normalized = pd.value_counts(votes, normalize=True)
print(result_normalized)
print(f"\nSum of proportions: {result_normalized.sum():.2f}")
# 3. Sort in ascending order
print("\n=== Sorted in ascending order ===")
print(pd.value_counts(votes, ascending=True))
Expected output:
=== 投票数据 === 0 A 1 B 2 A 3 C 4 A 5 B 值 'A' 出现 6 次,占 50% 值 'B' 出现 4 次,约 33.3% 值 'C' 出现 2 次,约 16.7% === 相对比例 === A 0.500000 B 0.333333 C 0.166667 dtype: float64 === 升序排列 === C 2 B 4 A 6 dtype: int64
Code explanation:
- Using
normalize=Truecan quickly calculate the proportion of each category. - All proportions add up to 1.0, making it convenient for proportion analysis.
- Using
ascending=Trueallows you to view the results in ascending order of frequency.
Example 3: Handling Numeric Data and Using the bins Parameter
For continuous numeric values, you can use thebinsparameter to perform binning statistics.
Example
import numpy as np
# 1. Create numeric data
scores = pd.Series([85, 90, 78, 92, 88, 76, 95, 82, 70, 89, 91, 77, 84, 86])
print("=== Student scores ===")
print(scores)
# 2. Without binning - count each unique value
print("\n=== Count each score ===")
print(pd.value_counts(scores))
# 3. Use the bins parameter for binning
print("\n=== Divided into 4 bins ===")
result_bins = pd.value_counts(scores, bins=4)
print(result_bins)
# 4. View the bin edges
print("\n=== Bin edges ===")
print(f"Min value: {scores.min()}, Max value: {scores.max()}")
Expected output:
=== 学生成绩 === [85, 90, 78, 92, 88, 76, 95, 82, 70, 89, 91, 77, 84, 86] === 统计每个分数 === 85 2 90 1 78 1 92 1 88 1 76 1 95 1 82 1 70 1 89 1 91 1 77 1 84 1 86 1 Name: count, dtype: int64 === 分成 4 个箱 === result_bins = pd.value_counts(scores, bins=4) (88.75, 95.0] 5 (82.5, 88.75] 4 (76.25, 82.5] 3 (69.97399999999999, 76.25] 2 Name: count, dtype: int64 === 分箱边界 === 最小值: 70, 最大值: 95
This is similar to a histogram, showing that the score distribution is mainly concentrated in the medium-to-high range. The bins parameter divides numeric data into discrete bins by range, and then counts the amount of data in each bin.
Example 4: Handling Missing Values
Using thedropnaparameter controls whether missing values are counted.
Example
import numpy as np
# 1. Data containing missing values
data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', np.nan, 'b', 'a'])
print("=== Data containing missing values ===")
print(data)
# 2. NaN is not included by default (dropna=True)
print("\n=== NaN not included by default ===")
print(pd.value_counts(data))
# 3. Include NaN (dropna=False)
print("\n=== Including NaN ===")
print(pd.value_counts(data, dropna=False))
# 4. Count missing values for numeric data
numeric_data = pd.Series([1, 2, np.nan, 3, 2, np.nan, 1, 4])
print("\n=== Numeric data containing NaN ===")
print(pd.value_counts(numeric_data, dropna=False))
Expected output:
=== 包含缺失值的数据 === 0 a 1 b 2 NaN 3 a 4 None 5 c 6 NaN 7 b 8 a dtype: object === 默认不包含 NaN === a 3 b 2 c 1 dtype: 64 === 包含 NaN (dropna=False) === a 3 b 2 NaN 2 # 2 个 NaN (2 个 np.nan) c 1 dtype: int64 === 数值型数据包含 NaN === 1.0 2 2.0 2 NaN 2 # 2 个 NaN 3.0 1 4.0 1 dtype: int64
Code explanation:
- By default (
dropna=True), NaN values are not counted. - Using
dropna=Falseallows you to view the number of missing values, which is useful in data quality analysis. - It also applies to numeric data.
Other ExtensionsTip:
pd.value_counts()It is one of the most important functions in exploratory data analysis, allowing you to quickly understand the distribution of the data. If you only need to obtain a list of unique values without frequencies, use thepd.unique()function.
Common Pandas Functions