Pandas Series.quantile() Function

Pandas 常用函数Common Pandas Functions


Series.quantile()It is a function in Pandas used to calculate quantiles of a Series. Quantiles are values that divide ordered data into several equal parts; common ones include quartiles (25%, 50%, 75%), the median (50%), and so on.

Quantiles are important indicators for describing data distribution. They help understand the distribution shape and identify outliers. They are widely used in statistical analysis, score ranking, income analysis, and other scenarios.


Basic Syntax and Parameters

quantile()It is a member function of the Series object, called directly using the dot operator.

Syntax Format

Series.quantile(q=0.5, interpolation='linear', numeric_only=True, closed='both')

Parameter Description

Parameter Type Description Default Value
q float or array-like The quantile value(s), ranging from 0 to 1. Can be a single value or a list of multiple values. 0.5
interpolation str The interpolation method used when the quantile lies between two values. Optional values: 'linear', 'lower', 'higher', 'nearest', 'midpoint'. 'linear'
numeric_only bool If True, only numeric data is calculated. True
closed str Used in DataFrame to determine the closure of the interval. Not commonly used in Series. 'both'

Return Value

    • Return Type:floatorSeries
    • DescriptionReturns the value at the specified quantile. If q is a single value, returns a float; if q is a list, returns a Series.

    Examples

    Let's thoroughly master, through a series of examples from simple to complex,Series.quantile()the usage of Series.quantile().

    Example 1: Basic Usage - Calculating the Median

    The median is the 50% quantile and is the most commonly used quantile.

    Example

    import pandas as pd

    # Create a Series containing student scores
    scores = pd.Series([65, 70, 72, 75, 78, 80, 82, 85, 88, 90, 92, 95])

    print("Student scores:")
    print(scores)
    print()

    # Calculate the median (50% quantile)
    median_score = scores.quantile(0.5)

    print(f"Median: {median_score}")
    print(f"Using median() function: {scores.median()}")
    print()
    print("Analysis: 50% of student scores are less than or equal to 80.")

    Output:

    学生成绩:
    0    65
    1    70
    2    72
    3    75
    4    75
    5    80
    6    80
    7    85
    8    88
    9    90
    10    92
    11    95
    dtype: int64
    
    中位数:80.0
    
    中位数(使用 median() 函数):80.0
    

    Code explanation:

    • quantile(0.5)Equivalent tomedian()。
    • The median divides the data into two halves; 50% of the data is less than or equal to the median.

    Example 2: Calculating Quartiles

    Quartiles divide the data into four equal parts: 25%, 50%, 75%.

    Example

    import pandas as pd

    # Create a Series containing employee income
    income = pd.Series([3000, 3500, 3800, 4000, 4200, 4500, 5000, 5500, 6000, 8000, 15000])

    print("Employee monthly income data (yuan):")
    print(income)
    print()

    # Calculate the three quartiles
    q1 = income.quantile(0.25)   # First quartile (25%)
    q2 = income.quantile(0.50)   # Second quartile / median (50%)
    q3 = income.quantile(0.75)   # Third quartile (75%)

    print(f"First quartile Q1 (25%): {q1} yuan")
    print(f"Second quartile Q2 (50%): {q2} yuan")
    print(f"Third quartile Q3 (75%): {q3} yuan")
    print()

    # Calculate the interquartile range (IQR)
    iqr = q3 - q1
    print(f"Interquartile range IQR: {iqr} yuan")
    print()

    print("Analysis:")
    print("- 25% of employees earn less than or equal to 4000 yuan")
    print("- 50% of employees earn less than or equal to 5000 yuan")
    print("- 75% of employees earn less than or equal to 6000 yuan")
    print("- The larger the IQR, the more dispersed the data")

    Output:

    员工月收入数据(元):
    1    3000
    2    低于 3800
    3    4000
    4    4200
    5    4500
    6    5000
    7    5500
    8    6000
    9    8000
    10   15000
    dtype: int64
    
    Q1(25% 分位):4000.0 元
    Q2(50% 分位):5000.0 元
    Q3(75% 分位):6000.0 元
    
    四分位距 IQR = 6000 - 4000 = 2000 元
    

    Example 3: Calculating Multiple Quantiles at Once

    You can calculate multiple quantiles at once.

    Example

    import pandas as pd

    # Create data
    data = pd.Series([10, 20, 30, 40, 50, 60, 70, 80, 90, 100])

    print("Data:")
    print(data)
    print()

    # Calculate multiple quantiles
    percentiles = data.quantile([0, 0.1, 0.25, 0.5, 0.75, 0.9, 1.0])

    print("Quantiles:")
    print(percentiles)
    print()

    # The quantile can also be specified as a percentage string (single value only)
    print(f"Using string '50%': {data.quantile('50%')}")
    print(f"Using string '0.5': {data.quantile(0.5)}")

    Output:

    数据:
    0     10
    1     20
    2     30
    3     40
    4     50
    量值:60
    5     70
    6     80
    7     90
    8    100
    dtype: int 参数
    
    多个分位数:
    0.00    10.0
    0.10    19.0
    0.25    32.5
    0.50    50.0
    0.75    67.5
    0.90    81.0
    1.00   100.0
    dtype: float64
    

    Example 4: The Role of the interpolation Parameter

    When the quantile lies between two values, different interpolation methods produce different results.

    Example

    import pandas as pd

    # Create a Series with 6 elements
    data = pd.Series([10, 20, 30, 40, 50, 60])

    print("Data:", data.values)
    print()

    # Calculate the 30% quantile (between 20 and 30)
    # position = (n-1) * q = 5 * 0.3 = 1.5
    print("30% quantile using different interpolation methods:")

    linear = data.quantile(0.3, interpolation='linear')
    print(f"linear (linear interpolation, default): {linear}")

    lower = data.quantile(0.3, interpolation='lower')
    print(f"lower (take the smaller value): {lower}")

    higher = data.quantile(0.3, interpolation='higher')
    print(f"higher (take the larger value): {higher}")

    nearest = data.quantile(0.3, interpolation='nearest')
    print(f"nearest (take the nearest value): {nearest}")

    midpoint = data.quantile(0.3, interpolation='midpoint')
    print(f"midpoint (take the midpoint): {midpoint}")

    Output:

    数据:[10, 20, 30, 40, 50, 60]
    位置计算:(n-1) * q = 5 * 0.3 = 1.5,表示在 20 和 30 之间
    
    不同插值方法的结果:
    - linear(线性插值,默认):23.0
    - lower(向下取整):20.0
    - higher(向上取整):30.0
    - nearest(最近邻):20.0
    - midpoint(中间值):25.0
    

    Code explanation:

    • interpolation='linear'Performs linear interpolation between the two values, 20 + (30-20)*0.5 = 23.
    • interpolation='lower'Takes the value with the smaller index, i.e., 20.
    • interpolation='higher'Takes the value with the larger index, i.e., 30.
    • interpolation='nearest'Takes the value closest to the quantile position.
    • interpolation='midpoint'Takes the midpoint of the two values, i.e., (20+30)/2 = 25.

    Example 5: Using Quantiles to Identify Outliers

    The interquartile range (IQR) is often used to identify outliers in data.

    Example

    import pandas as pd

    # Create data containing some outliers
    data = pd.Series([12, 15, 18, 20, 22, 25, 28, 30, 32, 150])

    print("Data (with outliers):")
    print(data)
    print()

    # Calculate quartiles
    q1 = data.quantile(0.25)
    q3 = data.quantile(0.75)
    iqr = q3 - q1

    # Calculate outlier boundaries
    lower_bound = q1 - 1.5 * iqr
    upper_bound = q3 + 1.5 * iqr

    print(f"Q1(25%):{q1}")
    print(f"Q3(75%):{q3}")
    print(f"IQR:{iqr}")
    print()
    print(f"Lower bound of normal values: {lower_bound}")
    print(f"Upper bound of normal values: {upper_bound}")
    print()

    # Identify outliers
    outliers = data[(data  upper_bound)]
    print(f"Outliers: {outliers.values}")
    print()

    print("Analysis: 150 is an obvious outlier, exceeding the upper bound.")

    Output:

    数据(含异常值):
    3     150
    4     22
    5     25
    6     28
    7     30
    8     32
    9    150
    dtype: int64
    
    Q1(25%):18.75
    Q3(75%):30.25
    IQR = 11.5
    
    正常值下界:1.5
    正常值上界:47.5
    异常值:150
    
    分析:150 明显超出了正常范围,是一个异常值。
    

    Notes

    • The quantile value ranges from 0 to 1.
    • By default, the linear interpolation method is used to calculate quantiles.
    • When q is a list, the return value is a Series rather than a single value.
    • Quantiles can be used to identify outliers in data; a common method is the IQR rule (1.5 times the interquartile range).
    • For large datasets, quantile calculation is very efficient.

    Summary

    Series.quantile()It is an important function for analyzing data distribution. Its main features include:

    • Supports calculating any quantile (between 0 and 1).
    • Can calculate multiple quantiles at once.
    • Provides multiple interpolation methods to meet different needs.
    • Very useful for identifying outliers.

    In practical data analysis, quantiles are often used to understand the data distribution shape, compare different datasets, identify outliers, and so on. When combined with box plots, they can display the data distribution more intuitively.

    Pandas 常用函数Common Pandas Functions

    Other Extensions