Pandas Series.median() Function
Series.median()Is a function in Pandas used to calculate the median (middle value) of a Series. The median is the value located in the middle after sorting the data. It is not affected by extreme values and is a robust indicator of the central tendency of data.
When there are outliers or skewed distributions in the data, the median can more accurately reflect the central position of the data than the mean. It is often used in scenarios such as income data analysis, housing price statistics, and grade analysis.
Basic Syntax and Parameters
median()Is a member function of the Series object, called directly via the dot operator.
Syntax Format
Series.median(axis=None, skipna=True, level=None, numeric_only=None, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| axis | int | Specify the axis. A Series has only one row of data, so this parameter is mainly for compatibility with DataFrame. | None |
| skipna | bool | If True, NaN values are skipped during calculation; if False, the result returns NaN when encountering NaN. | True |
| level | int or str | If the Series has a MultiIndex, specify the level to calculate. | None |
| numeric_only | bool | If True, only numeric data is calculated; otherwise, it attempts to convert to numeric. | False |
Return Value
- Return Type:
float - DescriptionReturns the median of all elements in the Series. If the number of elements is even, returns the average of the middle two elements.
Examples
Let's thoroughly master ... through a series of examples from simple to complex.Series.median()the usage.
Example 1: Basic Usage - Median of an Odd Number of Elements
For data with an odd number of elements, the median is the value located in the middle after sorting.
Example
# Create a Series containing employee incomes
# Simulate the monthly income of 7 employees in a department (unit: thousand yuan)
income = pd.Series([5, 6, 7, 8, 9, 10, 50])
# Calculate the median
median_income = income.median()
print("Employee monthly income (thousand yuan):")
print(income)
print()
print(f"Average income: {income.mean():.2f} thousand yuan")
print(f"Median income: {median_income} thousand yuan")
print()
print("Note: The average is greatly affected by the extreme value (50 thousand yuan), while the median better reflects the true level.")
Output:
员工月收入(千元): 0 5 1 6 2 7 3 8 4 9 5 10 6 50 dtype: int64 平均收入:13.57 千元 中位数收入:8.0 千元
Code Analysis:
- After sorting the data: [5, 6, 7, 8, 9, 10, 50]
- The middle position is the 4th element (index 3), with a value of 8.
- Because there is an extreme value (50), the average (13.57) is pulled up, while the median (8) better reflects the income level of most people.
Example 2: Median of an Even Number of Elements
For data with an even number of elements, the median is the average of the middle two elements.
Example
# Create a Series with 6 elements
# Simulate the exam scores of 6 students
scores = pd.Series([75, 82, 88, 92, 95, 100])
# Calculate the median
median_score = scores.median()
print("Student exam scores:")
print(scores)
print()
print(f"Average score: {scores.mean():.2f}")
print(f"Median score: {median_score}")
Output:
学生考试成绩: 0 75 1 82 2 88 3 92 4 95 5 100 dtype: int64 平均成绩:88.67 中位数成绩:90.0
Code Analysis:
- After sorting the data: [75, 82, 88, 92, 95, 100]
- The middle two elements are 88 and 92 (indexes 2 and 3).
- Median = (88 + 92) / 2 = 90.
Example 3: Handling Data Containing Missing Values
skipnaThe parameter determines how missing values are handled.
Example
import numpy as np
# Create a Series with missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40, np.nan, 50])
print("Data with missing values:")
print(data_with_nan)
print()
# By default skipna=True, skip NaN when calculating the median
median_skipna = data_with_nan.median()
print(f"Median when skipna=True (default): {median_skipna}")
# Set skipna=False
median_no_skipna = data_with_nan.median(skipna=False)
print(f"Median when skipna=False: {median_no_skipna}")
Output:
包含缺失值的数据: 0 10.0 1 20.0 2 NaN 3 30.0 4 40.0 5 NaN 6 50.0 dtype: float64 skipna=True(默认)时的中位数:30.0 skipna=False 时的中位数:nan
Code Analysis:
- Valid data: [10, 20, 30, 40, 50], 5 elements in total.
- After sorting, take the middle value: 30.
skipna=FalseWhen ..., as long as there is a NaN, it returns NaN.
Example 4: Application Scenarios Comparing Mean and Median
Demonstrates the advantage of the median when dealing with skewed data.
Example
# Simulate the monthly income data of 10 households in a residential community
# Most people have incomes between 3000-5000 yuan, but there are a few high-income individuals
income_data = pd.Series([3000, 3500, 3800, 4000, 4200, 4500, 4800, 5000, 8000, 50000])
print("Monthly income data of community residents (yuan):")
print(income_data)
print()
mean_income = income_data.mean()
median_income = income_data.median()
print(f"Average: {mean_income:.2f} yuan")
print(f"Median: {median_income} yuan")
print()
print("Analysis:")
print("The average (10380 yuan) is greatly pulled up by the extreme value (50000 yuan).")
print("The median (4350 yuan) better reflects the true income level of most households.")
print("This is why statistical departments more commonly use the median when publishing income data.")
Output:
小区住户月收入数据(元): 0 3000 1 3500 2 3800 3 4000 4 4200 5 4500 6 4800 7 5000 8 8000 9 50000 dtype: int64 平均值:10380.00 元 中位数:4350.0 元
Notes
- The median is insensitive to extreme values and is more robust than the mean.
- When the number of data points is even, the median is the average of the middle two numbers.
- When the data has a significantly skewed distribution, the median better reflects the central position of the data than the mean.
skipnaThe behavior of the parameter ismean()the same.
Summary
Series.median()It is an important function for describing the central tendency of data. Its main features include:
- Not affected by extreme values, the calculation result is robust.
- For an even number of elements, returns the average of the middle two elements.
- Especially useful in data analysis with skewed distributions such as income and housing prices.
- The syntax and parameters are
mean()exactly the same, easy to learn and use.
In practical data analysis, it is recommended to calculate both the mean and the median to more comprehensively understand the distribution characteristics of the data. When the difference between the two is large, it indicates that the data may have a skewed distribution or outliers.
Other Extensions
Pandas Common Functions