Pandas Series.sum() Function
Series.sum()It is a function in Pandas used to calculate the sum of all elements in a Series. It is one of the most commonly used statistical functions in data analysis and can quickly obtain the sum of numeric data.
Whether calculating total sales, total profit, or total quantity,sum()it can help you complete it quickly. It is especially suitable for handling financial data, statistical reports, sales data analysis, and similar scenarios.
Basic Syntax and Parameters
sum()It is a member function of the Series object and is called directly via the dot operator.
Syntax Format
Series.sum(axis=None, skipna=True, level=None, numeric_only=None, min_count=0, **kwargs)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| axis | int | Specifies the axis. A Series has only one row of data; this parameter is mainly for compatibility with DataFrame. | None |
| skipna | bool | If True, NaN values are skipped during calculation; if False, then the result returns NaN when NaN is encountered. | True |
| level | int or str | If the Series has a MultiIndex, specify the level to calculate. | None |
| numeric_only | bool | If True, only numeric data is calculated; otherwise, it attempts to convert to numeric values. | False |
| min_count | int | The minimum number of valid values required for calculation. If the number of valid values is less than this count, NaN is returned. | 0 |
Return Value
- Return Type: numeric (
int、floatornumpy.nan) - Description: Returns the sum of all elements in the Series. If all elements are NaN (and skipna=True), returns 0.
Examples
Let us, through a series of examples from simple to complex, thoroughly masterSeries.sum()the usage of it.
Example 1: Basic Usage - Calculate the Sum of a Numeric List
The most basic usage is to create a numeric Series and then callsum()to calculate the sum.
Example
# Create a Series containing sales amounts
# Simulate daily sales data for a week
sales_data = pd.Series([1200, 1500, 1800, 900, 2100, 1600, 1350])
# Calculate total sales
total_sales = sales_data.sum()
print("Daily sales:")
print(sales_data)
print()
print(f"Total sales: {total_sales}")
Output:
每日销售额: 0 1200 1 1500 2 1800 3 900 4 2100 5 1600 6 1350 dtype: int64 总销售额:10450
Code explanation:
- Created a Series containing 7 days of sales data, with an integer data type.
sum()It simply iterates over all elements and adds them up to get the total 10450.
Example 2: Handling Data Containing Missing Values
Missing values often occur in real data,skipnaand the parameter determines how to handle these missing values.
Example
import numpy as np
# Create a Series containing missing values
# Simulate missing sales data for certain dates
sales_with_nan = pd.Series([1200, 1500, np.nan, 1800, np.nan, 2100, 1600])
print("Sales data with missing values:")
print(sales_with_nan)
print()
# Default skipna=True, skip NaN and calculate the sum
total_skipna = sales_with_nan.sum()
print(f"Sum with skipna=True (default): {total_skipna}")
# Set skipna=False, return NaN when NaN is encountered
total_no_skipna = sales_with_nan.sum(skipna=False)
print(f"Sum with skipna=False: {total_no_skipna}")
Output:
包含缺失值的销售数据: 0 1200.0 1 1500.0 2 NaN 3 1800.0 4 NaN 5 2100.0 6 1600.0 dtype: float64 skipna=True(默认)时的总和:8200.0 skipna=False 时的总和:nan
Code explanation:
- When
skipna=TrueWhen (the default value),sum()it automatically skips NaN values and calculates the sum of only valid values. - Calculation process: 1200 + 1500 + 1800 + 2100 + 1600 = 8200
- When
skipna=FalseWhen skipna=False, as long as NaN exists, the result returns NaN. This is useful in scenarios where you need to clearly know data integrity.
Example 3: Using the min_count Parameter
min_countThe min_count parameter can set the minimum number of valid values required for calculation. This is useful when ensuring data sufficiency.
Example
import numpy as np
# Create a Series where most values are missing
sparse_data = pd.Series([np.nan, np.nan, 100, np.nan, np.nan])
print("Data with most values missing:")
print(sparse_data)
print()
# Default min_count=0, calculate as long as there are at least 0 valid values
result_default = sparse_data.sum()
print(f"Result with min_count=0 (default): {result_default}")
# Set min_count=3, require at least 3 valid values to calculate
result_min_count = sparse_data.sum(min_count=3)
print(f"Result with min_count=3: {result_min_count}")
# Set min_count=2, only need 2 valid values
result_min_count_2 = sparse_data.sum(min_count=2)
print(f"Result with min_count=2: {result_min_count_2}")
Output:
大部分缺失的数据: 0 NaN 1 NaN 2 100.0 3 NaN 4 NaN dtype: float64 min_count=0(默认)时的结果:100.0 min_count=3 时的结果:nan min_count=2 时的结果:100.0
Code explanation:
- This Series has only 1 valid value (100).
min_count=3This means at least 3 valid values are required for calculation, but there is actually only 1, so NaN is returned.min_count=2This means at least 2 valid values are required, but there is actually 1 valid value, which does not meet the condition, so NaN is returned.- This parameter is very useful in analysis scenarios with high data quality requirements, as it can help identify situations where valid data is insufficient.
Example 4: Application in Real Data Analysis
Combining real business scenarios, demonstratesum()its typical applications.
Example
# Create a simulated sales data DataFrame
sales_data = pd.DataFrame({
'Month': ['Jan', 'Feb', 'Mar', 'Apr', 'May', 'Jun'],
'East China Region': [12000, 15000, 18000, 16000, 14000, 17000],
'North China Region': [10000, 11000, 13000, 12500, 11500, 14000],
'South China Region': [15000, 17000, 19000, 18000, 16000, 20000]
})
print("Sales data by region for the first half of the year:")
print(sales_data)
print()
# Calculate the half-year total sales for each region
for region in ['East China Region', 'North China Region', 'South China Region']:
total = sales_data[region].sum()
print(f"{region} half-year total sales: {total} yuan")
print()
# Calculate the total sales for all regions
total_all = sales_data[['East China Region', 'North China Region', 'South China Region']].sum().sum()
print(f"Total half-year sales for all regions: {total_all} yuan")
Output:
上半年各区域销售数据: 月份 华东区 华北区 华南区 0 1月 12000 10000 15000 1 2月 15000 11000 17000 2 3月 18000 13000 19000 3 4月 16000 12500 18000 4 5月 14000 11500 16000 5 6月 17000 14000 20000 华东区半年总销售额:92000 元 华北区半年总销售额:72000 元 华南区半年总销售额:105000 元 所有区域半年总销售额:269000 元
Notes
sum()By default, NaN values are skipped, which is the expected behavior in most cases.- If you need to sum a Series containing non-numeric data, you should first clean the data or use the
numeric_only=Truenumeric_only parameter. min_countThe min_count parameter is very useful in data validation and quality checks, as it can help identify situations where valid data is insufficient.- For large datasets,
sum()sum() usually performs well because it uses NumPy's vectorized operations under the hood.
Summary
Series.sum()It is one of the most basic and commonly used statistical functions in Pandas. Its main features include:
- Simple and easy to use, called directly through the dot operator.
- Supports handling missing values via the
skipnaskipna parameter. - You can use the
min_countmin_count parameter to set the minimum requirement for valid values. - It uses NumPy's optimized implementation at the underlying level, resulting in high computational efficiency.
In real data analysis,sum()it is usually used together with other aggregation functions (such asmean()、count()mean(), etc.) to conduct comprehensive statistical analysis of data.
Common Pandas Functions