Pandas pd.qcut() Function

Pandas 通用函数Pandas Common Functions


pd.qcut()is a function in the Pandas library used forbinning by quantiles. It divides data into an equal number of samples, with each bin containing approximately the same number of data points.

andpd.cut()Unlike (equal-width binning),pd.qcut()it ensures that each interval contains an equal amount of data, making it very suitable for handling skewed distributions or cases where data needs to be divided according to its natural quantiles.

Word Meanings: qcutIn it, "q" represents "quantile" (Dividebitsnumber), meaning to "cut" data by quantiles.


Basic Syntax and Parameters

pd.qcut()It is a top-level function of the Pandas library, used for discretization based on quantiles.

Syntax Format

pd.qcut(x, q, labels=None, retbins=False, precision=3, duplicates='raise')

Parameter Description

  • Parameter: x
    • Type: Array-like object, such as a list, Series, array, etc.
    • Description: Continuous numeric data to be binned.
  • Parameter: q
    • Type: Integer or quantile array.
    • Description: The number of bins or the specific quantiles. If it is an integer n, it means dividing the data into n groups of equal size (each group contains about 1/n of the data). It can also be an array of quantiles (e.g., [0, 0.25, 0.5, 0.75, 1.0]).
  • Parameter: labels
    • Type: Array or None.
    • Description: Specify custom labels for each interval. If not specified, quantile notation (such as "(0, 25]") is used.
  • Parameter: retbins
    • Type: Boolean value.
    • Description: IfTrue, the returned result includes the bin boundary values. Default isFalse。
  • Parameter: duplicates
    • Type: String ('raise' or 'drop').
    • Description: If the bin edges have duplicates (resulting in fewer intervals than the specified q value),'raise'raises an exception,'drop'deletes duplicate edges. Default is'raise'。

Function Description

  • Return Value: Returns a Series of Categorical type, each value corresponding to its quantile group.
  • Effect: Maps continuous values into discrete categories by quantiles, each category containing approximately the same number of samples.

Examples

Let us thoroughly master, through a series of examples from simple to complex,pd.qcut()the usage of.

Example 1: Basic Usage - Binning by Quartile

Example

import pandas as pd
import numpy as np

# 1. Create skewed numeric data (most values are small)
np.random.seed(42)
data = np.random.exponential(scale=2, size=100)  # Exponential distribution data

# Take a subset of the data for readability
scores = pd.Series(data[:20])

print("=== Original data ===")
print(scores.sort_values().values)

# 2. Use pd.qcut() to divide the data into 4 quantile groups (quartile)
result = pd.qcut(scores, q=4)
print("\n=== pd.qcut(scores, q=4) quartile binning ===)
print(result.value_counts().sort_index())

Expected output:

=== 原始数据 ===
[ 0.078  0.29   0.68   0.91   1.03  1.06  1.23  1.3   1.5   1.69
  1.86  2.01  2.12  2.3   2.32  2.42  3.03  3.39  5.29  7.39]

=== pd.qcut(scores, q=4) 四分位分箱 ===
(0.0783, 1.547]     5
(1.547, 2.315]     5
(2.315, 3.151]     5
(3.151, 7.393]     5
dtype: int64

Code explanation:

  1. For exponentially distributed data, if equal-width binning is used, most data will concentrate in one bin.
  2. pd.qcut()It ensures that each bin contains approximately the same number of samples (here 5, exactly 20/4).
  3. The quantile points are calculated based on the actual distribution of the data, rather than being evenly divided.

Example 2: Specifying Custom Quantiles

Through theqparameter, you can specify arbitrary quantiles, such as deciles, percentiles, etc.

Example

import pandas as pd
import numpy as np

# 1. Create data
df = pd.DataFrame({
    'name': ['Student_' + str(i) for i in range(1, 11)],
    'score': [65, 78, 82, 55, 90, 72, 68, 85, 45, 95]
})

print("=== Student score data ===")
print(df)

# 2. Bin by custom quantiles [0, 0.3, 0.7, 1.0] representing 30%, 70%, 100%
result = pd.qcut(df['score'], q=[0, 0.3, 0.7, 1.0])
print("\n=== Binning by 30% and 70% quantiles ===)
print(result)

# 3. Add custom labels
labels = ['Pass', 'Good', 'Excellent']
result_labeled = pd.qcut(df['score'], q=3, labels=labels)
print("\n=== Custom labels added ===)
print(result_labeled)

Expected output:

=== 学生成绩数据 ===
      name  score
0  Student_1     65
1  Student_2     78
2  Student_3     82
3  Student_4     55
4  Student_5     90
5  Student_6     72
6  Student_7     68
7  Student_8     85
8  Student_9     45
9  Student_10    95

=== 按 30% 和 70% 分位数分箱 ===
0      (44.999, 65.8]
1      (65.8, 86.4]
2      (65.8, 86.4]
3      (44.999, 65.8]
4      (86.4, 95.0]
5      (65.8, 86. 86.4]
6      (0.0, 44.999]
7      (44.999, 65.8]
8      (0.0, 44.999]
9      (65.8, 86.4]
10     (65.8, 86.4]

Code explanation:

  • q=[0, 0.3, 0.7, 1.0]Three intervals are defined: 0-30% (low scores), 30%-70% (medium), and 70%-100% (high scores).
  • Using thelabelsparameter, you can name each interval instead of using the default interval notation.

Example 3: Comparison of cut vs qcut

Through comparison, you can more clearly understandcutandqcutthe difference.

Example

import pandas as pd
import numpy as np

# 1. Create skewed data (most values are small)
data = pd.Series([1, 1, 1, 1, 1, 1, 1, 1, 1, 10, 20, 30, 40, 50])

print("=== Original data ===")
print(data.values)

# 2. Equal-width binning (pd.cut)
print("\n=== pd.cut equal-width binning ===)
cut_result = pd.cut(data, bins=4)
print(cut_result.value_counts().sort_index())

# 3. Equal-frequency binning (pd.qcut)
print("\n=== pd.qcut equal-frequency binning ===)
qcut_result = pd.qcut(data, q=4)
print(qcut_result.value_counts().sort_index())

Expected output:

=== 原始数据 ===
[ 1  1  1  1  1  1  1  1  1 10 20 30 40 50]

=== pd.cut 等宽分箱 ===
(0.975, 13.25]    10   # 10 个值在这个区间
(13.25, 25.5]      1
(25.5, 37.75]      1
(37.75, 50.0]      2   # 2 个值在这个区间
dtype: 50

=== pd.qcut 等量分箱 ===
(0.999, 7.75]      4   # 约 1/4 的数据
(7.75, 25.0]       4
(25.0, 42.5]       4
(10.0, 50.0]       4   # 约 1/4 的数据

Code explanation:

  • pd.cut()The data range is evenly divided into 4 parts, but because the data is highly skewed, most values (nine 1s) fall into the first interval.
  • pd.qcut()It ensures each interval contains the same number of data points (4), even if the interval widths differ greatly.

Tip: pd.cut()Suitable for cases where data distribution is uniform or fixed-width intervals are needed.pd.qcut()Suitable for cases where data distribution is uneven or division by rank/percentage is required.

Pandas 常用函数Pandas Common Functions

Other Extensions