Pandas pd.value_counts() Function

Pandas 通用函数Common Pandas Functions


pd.value_counts()is a function in the Pandas library used forcounting the frequency of each value. It calculates the number of occurrences of each unique value in the array and returns them sorted by frequency in descending order.

This is one of the most commonly used functions in categorical data analysis, allowing you to quickly understand the distribution of the data, such as counting voting results, product category distributions, etc.

Word Meaning: value_countsIt means "count values", i.e., counting the number of occurrences of each value.


Basic Syntax and Parameters

pd.value_counts()It is a top-level function of the Pandas library, used to calculate the frequency of each unique value.

Syntax Format

pd.value_counts(values, sort=True, ascending=False, normalize=False, bins=None, dropna=True)

Parameter Description

  • Parameter: values
    • Type: Series, array-like object.
    • Description: The data whose frequencies are to be counted. Usually a Series.
  • Parameter: sort
    • Type: Boolean.
    • Whether to sort by frequency. Defaults toTrue(descending order by frequency).
  • Parameter: ascending
    • Type: Boolean.
    • IfTrue, sort by frequency in ascending order; ifFalse(default), sort by frequency in descending order.
  • Parameter: normalize
    • Type: Boolean.
    • IfTrue, return the proportion of each value (between 0 and 1) instead of the absolute frequency. Defaults toFalse。
  • Parameter: bins
    • Type: Integer or None.
    • If an integer is specified, the numeric data is binned and frequencies are counted (similar to a histogram). Not applicable to categorical data.
  • Parameter: dropna
    • Type: Boolean.
    • Whether to include NaN in the statistics results. Defaults toTrue(not included).
  • Function Description

    • Return Value: Returns a Series where the index is the unique values and the values are the frequencies (or proportions).
    • Effect: Counts and displays the number of occurrences of each unique value.

    Examples

    Let's thoroughly masterpd.value_counts()the usage of.

    Example 1: Basic Usage - Counting Frequencies of Categorical Data

    Example

    import pandas as pd
    import numpy as np

    # 1. Create a Series containing duplicate values
    colors = pd.Series(['red', 'blue', 'green', 'red', 'blue', 'yellow', 'red', 'blue', 'blue'])

    print("=== Original Series ===")
    print(colors)

    # 2. Use pd.value_counts() to count frequencies
    result = pd.value_counts(colors)
    print("\n=== pd.value_counts() frequency statistics ===")
    print(result)

    Expected output:

    === 原始 Series ===
    0       red
    1      blue
    2     green
    3       red
    4     blue
    5     yellow
    6       red
    7      blue
    8       blue
    dtype: object
    
    === pd.value_counts() 频次统计 ===
    blue      4
    red       3
    green     1
    yellow    1
    dtype: int64
    

    Code explanation:

    1. The results are sorted by frequency in descending order; blue appears the most (4 times), red comes next (3 times), and green and yellow each appear 1 time.
    2. What is returned is a Series, with the color values as the index and the counts as the values.

    Example 2: Using the normalize Parameter to Calculate Proportions

    normalize=TrueFrequencies can be converted into proportions, making it easier to view the relative distribution of the data.

    Example

    import pandas as pd
    import numpy as np

    # Create voting data
    votes = pd.Series(['A', 'B', 'A', 'C', 'A', 'B', 'A', 'A', 'C', 'B', 'A', 'B'])

    print("=== Voting data ===")
    print(votes)

    # 1. Absolute frequency
    print("\n=== Absolute frequency ===")
    print(pd.value_counts(votes))

    # 2. Relative proportions (normalize=True)
    print("\n=== Relative proportions ===")
    result_normalized = pd.value_counts(votes, normalize=True)
    print(result_normalized)
    print(f"\nSum of proportions: {result_normalized.sum():.2f}")

    # 3. Sort in ascending order
    print("\n=== Sorted in ascending order ===")
    print(pd.value_counts(votes, ascending=True))

    Expected output:

    === 投票数据 ===
    0        A
    1        B
    2        A
    3        C
    4        A
    5        B
    值 'A' 出现 6 次,占 50%
    值 'B' 出现 4 次,约 33.3%
    值 'C' 出现 2 次,约 16.7%
    
    === 相对比例 ===
    A    0.500000
    B    0.333333
    C    0.166667
    dtype: float64
    
    === 升序排列 ===
    C    2
    B    4
    A    6
    dtype: int64
    

    Code explanation:

    • Usingnormalize=Truecan quickly calculate the proportion of each category.
    • All proportions add up to 1.0, making it convenient for proportion analysis.
    • Usingascending=Trueallows you to view the results in ascending order of frequency.

    Example 3: Handling Numeric Data and Using the bins Parameter

    For continuous numeric values, you can use thebinsparameter to perform binning statistics.

    Example

    import pandas as pd
    import numpy as np

    # 1. Create numeric data
    scores = pd.Series([85, 90, 78, 92, 88, 76, 95, 82, 70, 89, 91, 77, 84, 86])

    print("=== Student scores ===")
    print(scores)

    # 2. Without binning - count each unique value
    print("\n=== Count each score ===")
    print(pd.value_counts(scores))

    # 3. Use the bins parameter for binning
    print("\n=== Divided into 4 bins ===")
    result_bins = pd.value_counts(scores, bins=4)
    print(result_bins)

    # 4. View the bin edges
    print("\n=== Bin edges ===")
    print(f"Min value: {scores.min()}, Max value: {scores.max()}")

    Expected output:

    === 学生成绩 ===
    [85, 90, 78, 92, 88, 76, 95, 82, 70, 89, 91, 77, 84, 86]
    
    === 统计每个分数 ===
    85    2
    90    1
    78    1
    92    1
    88    1
    76    1
    95    1
    82    1
    70    1
    89    1
    91    1
    77    1
    84    1
    86    1
    Name: count, dtype: int64
    
    === 分成 4 个箱 ===
      result_bins = pd.value_counts(scores, bins=4)
    (88.75, 95.0]                 5
    (82.5, 88.75]                 4
    (76.25, 82.5]                 3
    (69.97399999999999, 76.25]    2
    Name: count, dtype: int64
    
    === 分箱边界 ===
    最小值: 70, 最大值: 95
    

    This is similar to a histogram, showing that the score distribution is mainly concentrated in the medium-to-high range. The bins parameter divides numeric data into discrete bins by range, and then counts the amount of data in each bin.

    Example 4: Handling Missing Values

    Using thedropnaparameter controls whether missing values are counted.

    Example

    import pandas as pd
    import numpy as np

    # 1. Data containing missing values
    data = pd.Series(['a', 'b', np.nan, 'a', None, 'c', np.nan, 'b', 'a'])

    print("=== Data containing missing values ===")
    print(data)

    # 2. NaN is not included by default (dropna=True)
    print("\n=== NaN not included by default ===")
    print(pd.value_counts(data))

    # 3. Include NaN (dropna=False)
    print("\n=== Including NaN ===")
    print(pd.value_counts(data, dropna=False))

    # 4. Count missing values for numeric data
    numeric_data = pd.Series([1, 2, np.nan, 3, 2, np.nan, 1, 4])
    print("\n=== Numeric data containing NaN ===")
    print(pd.value_counts(numeric_data, dropna=False))

    Expected output:

    === 包含缺失值的数据 ===
    0       a
    1       b
    2     NaN
    3       a
    4    None
    5       c
    6     NaN
    7       b
    8       a
    dtype: object
    
    === 默认不包含 NaN ===
    a    3
    b    2
    c    1
    dtype: 64
    
    === 包含 NaN (dropna=False) ===
    a       3
    b       2
    NaN     2  # 2 个 NaN (2 个 np.nan)
    c       1
    dtype: int64
    
    === 数值型数据包含 NaN ===
    1.0    2
    2.0    2
    NaN    2  # 2 个 NaN
    3.0    1
    4.0    1
    dtype: int64
    

    Code explanation:

    • By default (dropna=True), NaN values are not counted.
    • Usingdropna=Falseallows you to view the number of missing values, which is useful in data quality analysis.
    • It also applies to numeric data.

    Tip: pd.value_counts()It is one of the most important functions in exploratory data analysis, allowing you to quickly understand the distribution of the data. If you only need to obtain a list of unique values without frequencies, use thepd.unique()function.

    Pandas 常用函数Common Pandas Functions

    Other Extensions