Pandas pd.cut() Function

Pandas 通用函数Pandas Common Functions


pd.cut()is used in the Pandas library tobin continuous data (Binning)function. It divides numerical data into discrete categories according to specified intervals, and is very suitable for grouped statistics and visualization in data analysis.

Binning operations are very common in data analysis, such as dividing ages into "teenager", "middle-aged", "elderly", or dividing scores into "fail", "pass", "good", "excellent", etc.

Word Meaning: cutIt means "cut". Here it refers to "cutting" a continuous data interval into multiple discrete intervals.


Basic Syntax and Parameters

pd.cut()is a top-level function of the Pandas library, used to discretize continuous variables.

Syntax Format

pd.cut(x, bins, right=True, labels=None, retbins=False, precision=3, include_lowest=False)

Parameter Description

  • Parameter: x
    • Type: Array-like object, such as list, Series, array, etc.
    • Description: The continuous numerical data to be binned.
  • Parameter: bins
    • Type: Integer, scalar (interval boundaries), or IntervalIndex.
    • Description: The binning method. It can be an integer (indicating how many equal-width bins to divide into) or a list (specifying specific interval boundaries).
  • Parameter: right
    • Type: Boolean value.
    • Description: Whether to use right-closed intervals. Default isTrue(i.e., in the form [a, b]). If set toFalse, then it is a left-closed right-open interval (a, b].
  • Parameter: labels
    • Type: Array or None.
    • Description: Specify custom labels for each interval. If not specified, interval representations such as "(0, 10]" are used.
  • Parameter: retbins
    • Type: Boolean value.
    • Description: IfTrue, the return result will contain the bin boundaries. Default isFalse。
  • Parameter: include_lowest
    • Type: Boolean value.
    • Description: IfTrue, the first interval contains the left boundary. Default isFalse。

Function Description

  • Return Value: Returns a Series of Categorical type, where each value corresponds to the interval to which it belongs.
  • Effect: Maps continuous numerical values to discrete categories (intervals).

Example

Let's thoroughly master, through a series of examples from simple to complex,pd.cut()the usage.

Example 1: Basic Usage — Equal-Width Binning

Example

import pandas as pd
import numpy as np

# 1. Create numerical data
scores = pd.Series([55, 70, 82, 45, 90, 65, 78, 88, 92, 58])

print("=== Original scores ===")
print(scores)

# 2. Divide the scores into 3 equal-width intervals
result = pd.cut(scores, bins=3)
print("\n=== pd.cut(scores, bins=3) equal-width binning ===)
print(result)

# 3. View the bin boundaries
result_bins = pd.cut(scores, bins=3, retbins=True)
print("\n=== Return bin boundaries ===)
print(f"Bin boundaries: {result_bins)

Expected output:

=== 原始分数 ===
0    55
1    70
2    82
3    45
4    90
5    65
6    78
7    88
8    92
9    58

=== pd.cut(scores, bins=3) 等宽分箱 ===
0    (44.667, 60.667]
1    (60.667, 76.333]
2    (76.333, 92.0]
3    (44.667, 60.667]
4    (76.333, 92.0]
5    (60.667, 76.333]
6    (60.667, 76.333]
7    (76.333, 92.0]
8    (76.333, 92.0]
9    (44.667, 60.667]
dtype: category
Categories (3, interval[float64]): [(44.667, 60.667] < (60.667, 76.333] < (76.333, 92.0]

=== 返回分箱边界 ===
分箱边界: [44.667 60.667 76.333 92.   ]

Code explanation:

  1. bins=3Indicates that the data is automatically divided into 3 equal-width intervals.
  2. Pandas automatically calculates the interval boundaries based on the minimum and maximum values of the data.
  3. The return value is of Categorical type, and each value is an Interval object.
  4. By default, right-closed intervals are used, i.e., in the form (a, b].

Example 2: Custom Interval Boundaries

In practical applications, we often need to customize bin boundaries based on business requirements.

Example

import pandas as pd
import numpy as np

# Create score data
scores = pd.Series([55, 70, 82, 45, 90, 65, 78, 88, 92, 58, 73, 67, 81, 49, 95])

# Custom bin boundaries: Fail, Pass, Good, Excellent
bins = [0, 60, 75, 90, 100]
labels = ['Fail', 'Pass', 'Good', 'Excellent']

result = pd.cut(scores, bins=bins, labels=labels, include_lowest=True)

print("=== Original scores ===")
print(scores.values)
print("\n=== Binning result ===)
print(result)

# Count the number of people in each grade
print("\n=== Score distribution statistics ===)
print(result.value_counts().sort_index())

Expected output:

=== 原始成绩
[55 70 82 45 90 65 78 88 92 58 73 67 81 49 95]

=== 分箱结果
0     及格
1     良好
2     良好
3    不及格
4     优秀
...
12    良好
13    不及格
14     优秀
dtype: category
Categories (4, interval[int64]): [不及格 < 及格 < 良好 < 优秀]

=== 成绩分布统计
不及格     3
及格       4
良好       5
优秀       3
dtype: int64

Code explanation:

  • bins=[0, 60, 75, 90, 100]Defined 4 intervals: [0-60], (60-75], (75-90], (90-100]
  • labelsThe labels parameter specifies intuitive Chinese labels for each interval
  • include_lowest=TrueEnsure that the first interval contains the left boundary 0
  • value_counts()Makes it easy to count the number of people in each grade

Example 3: Using Right-Closed/Left-Closed Intervals

Example

import pandas as pd
import numpy as np

# Test data
data = pd.Series([10, 20, 30, 40, 50])

# Default right-closed interval [a, b]
result1 = pd.cut(data, bins=4)
print("=== Default right-closed interval [a, b] ===")
print(result1)

# Left-closed right-open interval (a, b]
result2 = pd.cut(data, bins=4, right=False)
print("\n=== Left-closed right-open interval (a, b] ===)
print(result2)

Expected output:

=== 默认右闭区间 [a, b]
0     (7.5, 17.5]
1    (17.5, 27.5]
2    (27.5, 37.5]
3    (37.5, 47.5]
4    (47.5, 57.5]
dtype: category

=== 左闭右开区间 (a, b]
0    [10, 17.5)
1    [17.5, 27.5)
2    [27.5, 37.5)
3    [37.5, 47.5)
4    [47.5, 57.5)
dtype: category

Code explanation:

  • By default, the right-closed interval is used[a, b], i.e., includes the right boundary
  • Setright=Falseto change to a left-closed right-open interval[a, b)
  • The open/closed mode of the interval affects the classification of boundary values.

Notes

Important notes:

  • If there are values in the data that exceedbinsthe specified range, NaN will be producedNaN
  • binsboundaries must be monotonically increasing
  • When using custom labels, the number of labels must be one less than the number of boundaries.
  • The binned result is of Categorical type and can be processed using string methods.

Python math 模块Pandas Common Functions

Other Extensions