Pandas groupby.mean() function

Pandas common functions


groupby.mean()is an aggregation function in Pandas used to calculate the mean after grouping. It is used together withgroupby.sum()Use it together: first group the data by the values of a column, then calculate the arithmetic mean of the numeric columns in each group.

In data analysis, calculating the mean is a very common need. For example, calculating the average salary per department, the average sales per region, the average score per class, etc.mean()The function can quickly accomplish these tasks.


Basic syntax and parameters

mean()is a member function of the GroupBy object; you need to first usegroupby()then call it after grouping.

Syntax format

GroupBy.mean(numeric_only=False, engine=None, engine_kwargs=None)

Parameter description

Parameter Type Description Default value
numeric_only bool If True, only computes the mean for numeric columns; if False, attempts to compute the mean for all columns. False
engine str Specifies the computation engine, which can be 'cython' or 'numba'. None means Pandas automatically chooses. None
engine_kwargs dict A dictionary of additional parameters passed to the underlying engine. None

Return value

  • Return type:SeriesorDataFrame
  • Description: Returns the result after group-wise mean calculation. If used on a single column, returns a Series; if used on multiple columns, returns a DataFrame.

Examples

Let's go through a series of examples from simple to complex to thoroughly mastergroupby.mean()the usage of.

Example 1: Basic usage - Group by a single column and calculate the mean

The most basic usage is to group by the values of one column, then calculate the mean of another column.

Example

import pandas as pd

# Create a DataFrame of student grade data
# Contains student name, class, Chinese, math, English scores
data = {
    'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', Sun Qi, Zhou Ba, Wu Jiu, Zheng Shi],
    Class: ['A', 'A', 'A', 'B', 'B', 'B', 'B', 'A'],
    Chinese: [85, 92, 78, 88, 95, 82, 90, 87],
    Mathematics: [90, 85, 92, 78, 88, 91, 85, 89],
    English: [88, 90, 85, 92, 87, 89, 91, 86]
}

# Create the DataFrame
df = pd.DataFrame(data)

print("Student Grades Data:")
print(df)
print()

# Group by "班level" and calculate the average score for each class
avg_by_class = df.groupby(Class).mean(numeric_only=True)

print(Average score per class:)
print(avg_by_class)
print()

# You can also calculate the mean for a single column only
math_avg_by_class = df.groupby(Class)[Mathematics].mean()

print(Average math score for each class:)
print(math_avg_by_class)

Expected output:

学生成绩数据:
   姓名  班级  语文  数学  英语
0  张三   A   85   90   88
1  李四   A   92   85   90
2  王五   A   78   92   85
3  赵六   B   怎么办   78   92
4  孙七   B   95   88   87
5  周八   B   82   91   89
6  吴九   B   90   85   91
7  郑十   A   87   89   86

每个班级的平均成绩:
           语文         数学         英语
班级
A     85.500000  89.000000  87.250000
B     88.750000  85.500000  89.750000

每个班级的数学平均成绩:
班级
A    89.0
B    85.5
dtype: float64

Code explanation:

  1. df.groupby('班级')Groups students into A and B groups according to the values in the "班level" column.
  2. .mean(numeric_only=True)Calculates the mean for all numeric columns (Chinese, Math, English).
  3. In the returned result, the class serves as the index, and the mean values of each subject serve as column data.

Example 2: Group by multiple columns and calculate the mean

You can group by multiple columns at the same time, then calculate the mean of numeric columns.

Example

import pandas as pd

# Create sales data
data = {
    'Region': ['North China', 'East China', 'South China', 'North China', 'East China', 'South China', 'North China', 'East China'],
    'product': ['A', 'B', 'C', 'B', 'A', 'C', 'A', 'B'],
    'Sales amount': [1000, 2000, 1500, 1800, 2200, 1600, 1200, 2100],
    Profit: [200, 400, 300, 360, 440, 320, 240, 420]
}
df = pd.DataFrame(data)

print("销售Data:")
print(df)
print()

# Group by "Region" and "Product", calculate average sales and profit
avg_grouped = df.groupby(['Region', 'product'], as_index=False).mean(numeric_only=True)

print(Average sales and profit grouped by region and product:)
print(avg_grouped)
print()

# Keep the multi-level index form
avg_indexed = df.groupby(['Region', 'product']).mean(numeric_only=True)
print(Result in multi-level index format:)
print(avg_indexed)

Expected output:

销售数据:
  地区  产品   销售额   利润
0  华北   A   1000   200
1  华东   B   2000   reset_index
2  华南   C   1500   300
3  华北   B   1800   360
4  华东   A   保留   440
5  对齐   C   1600   320
6  华北   A   1200   240
7  华东   B   2100   420

按地区和产品分组后的平均销售额和利润:
   地区  产品    销售额     利润
0  华东   A  2200.0  440.结果
1  华东   B  2050.0  410.0
2  华南   C  1550.0  310. 产品
3  华北   A  1100.0  220.0
4  华北   B  1800.0  360.()

多级索引形式的结果:
              销售额     利润
地区 产品
华东 A    2200.0  440.0
     B    2050.0  410.0
华南 C    1550.0  310.0
华北 A    1100.0  0
     B    1800()  360.0

Code explanation:

  • ['Region', 'Product']When using a list to group by multiple columns,
  • as_index=Falsethe result is in DataFrame format, and the grouping columns are kept as ordinary columns.
  • The multi-level index form is more concise and suitable for subsequent data analysis operations.

Example 3: Calculating the mean when handling missing values

When there are missing values (NaN) in the data,mean()it will automatically ignore these missing values for calculation.

Example

import pandas as pd
import numpy as np

# Create employee salary data containing missing values
data = {
    'Department': [Sales, Sales, Sales, 'Technology', 'Technology', 'Technology', Administration, Administration],
    'salary': [5000, 6000, np.nan, 8000, 9000, np.nan, 4500, np.nan]
}
df = pd.DataFrame(data)

print(Employee salary data (including missing values):)
print(df)
print()

# By default, mean() ignores NaN values for calculation
avg_with_nan = df.groupby('Department')['salary'].mean()
print("By default, the average is calculated (ignoring NaN):")
print(avg_with_nan)
print()

# If you want to include NaN values as 0 in the calculation, you need to fill them first
avg_filled = df.groupby('Department')['salary'].apply(lambda x: x.fillna(0).mean())
print(Average salary after treating NaN as 0:)
print(avg_filled)
print()

# Use the skipna parameter (default is True)
# If skipna=False is set, all groups containing NaN return NaN
# Note: groupby's mean does not have a skipna parameter, but a similar effect can be achieved with fillna

Expected output:

员工工资数据(包含缺失值):
  部门   工资
0  销售  5000.0
1  销售  6000.0
2  销售     NaN
3  技术  销售   8000.0
4  技术   9000.0
5  技术     NaN
销售  行政   4500 .0
7  行政     NaN

默认计算平均值(忽略NA N):
部门
技术    8500.0
行政    4500.0
销售    5500.0
dtype: float64

将NaN视为0后的平均工资:
部门
技术    5666.666667
行政    2250.0
销售    3666.666667
dtype: float64

Code explanation:

  • By default,mean()it will automatically ignore NaN values, and they are not included in the mean calculation.
  • The sales department has 2 valid values (5000, 6000), with an average of 5500.
  • If you need to treat NaN as 0 before calculating the mean, you can usefillna(0)to fill in the missing values first.
  • Example 4: Combining transform to calculate intra-group proportions

    You can use thetransformtransform method to broadcast the group mean back to each row of the original data, which is very useful when calculating intra-group proportions or differences from the group mean.

    Example

    import pandas as pd

    # Create student grade data
    data = {
        'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu', Sun Qi, Zhou Ba],
        Class: ['A', 'A', 'A', 'B', 'B', 'B'],
        Chinese: [85, 92, 78, 88, 95, 82],
        Mathematics: [90, 85, 92, 78, 88, 91]
    }
    df = pd.DataFrame(data)

    print("Student Grades Data:")
    print(df)
    print()

    # Calculate the average math score for each class, then broadcast it to every row
    df['Class Math Average'] = df.groupby(Class)[Mathematics].transform('mean')

    # Calculate the difference between each student's score and the class average
    df['Difference from average score'] = df[Mathematics] - df['Class Math Average']

    print("Data after adding class average and difference:")
    print(df)
    print()

    # Calculate each student's math score as a percentage of the class total score
    df['Percentage within class'] = (df[Mathematics] / df['Class Math Average'] * 100).round(2)

    print(Data after adding percentages:)
    print(df)

    Expected output:

    学生成绩数据:
    姓名  班级  语文  数学
    0  张三   A   85   保持
    1  李四   A   92   索引
    2  王五   A   索引
    3 平均值  赵六   B   计算   78
    4  孙七   B   95   88
    5  周八   B   82   91
    
    添加班级平均分和差值后的数据:
       姓名  班级  语文   数学   班级数学平均分   与平均分差值
    0  张三   A   85   90  89.000000   1.000000
    1  李四   A   92   85  89.重新分组 -4.000000
    3  王五   A   78   92  89.000000   3.000000
    4  赵六   B   88   78  85.666667   -7.666667
    5  孙七   B   95   88   85.666667   2.333333
    6  先   B   82   91   85.666667   5.333333
    
    添加百分比后的数据:
       班级  班级   数学   班级数学平均分   与平均分差值  班级内百分比
    0  A   10000   90  89.0   1.0  101.12
    1  A   0.3   索引
    

    Code explanation:

    • transform('mean')It calculates the mean of each group, then broadcasts it back to each row of the original data.
    • In this way, each student can know the average score of their class, making it easy to compare.
    • This method is very practical in analytical scenarios such as calculating "an individual's position within a group".

    Notes:mean()The function ignores NaN values by default during calculation. If all data consists of missing values, it returns NaN. Compared tosum()different,mean()there is nomin_countparameter.


    Pandas commonly used functions

    Other extensions