Pandas Series.cumsum() Function

Pandas 常用函数Common Pandas Functions


Series.cumsum()is a Pandas function used to calculate the cumulative sum of a Series. The cumulative sum means that starting from the first element, the value at each position equals the sum of all elements before that position (including the current element).

Cumulative sums are often used in financial data analysis (such as calculating cumulative returns), sales analysis (such as calculating cumulative sales), time series analysis, and other scenarios.


Basic Syntax and Parameters

cumsum()is a member function of the Series object, called directly via the dot operator.

Syntax

Series.cumsum(axis=None, skipna=True, dtype=None, out=None, **kwargs)

Parameter Description

Parameter Type Description Default Value
axis int Specifies the axis; a Series has only one column, so this parameter is mainly for compatibility with DataFrame. None
skipna bool If True, NaN values are skipped during calculation; if False, encountering NaN will break the cumulative sum. True
dtype dtype Specifies the output data type. None
out ndarray The array used to store the result; usually does not need to be set. None

Return Value

  • Return Type:Series
  • Description: Returns a new Series, where each position's value is the cumulative sum.

Examples

Through a series of examples from simple to complex, let us thoroughly masterSeries.cumsum()the usage.

Example 1: Basic Usage - Calculating Cumulative Sales

The cumulative sum is the total of all values up to the current position.

Example

import pandas as pd

# Create a Series containing daily sales
daily_sales = pd.Series([1000, 1500, 1200, 1800, 2000, 900, 1600])

print("Daily sales (yuan):")
print(daily_sales)
print()

# Calculate cumulative sales
cumulative_sales = daily_sales.cumsum()

print("Cumulative sales (yuan):")
print(cumulative_sales)
print()

print("Analysis:")
print("- Day 1 cumulative: 1000")
print("- Day 2 cumulative: 1000 + 1500 = 2500")
print("- Day 3 cumulative: 1000 + 1500 + 1200 = 3700")
print("And so on...")

Output:

每日销售额(元):
0    1000
1    1500
2    1200
3    1800
4    2000
5     900
6    1600
dtype: int64

累计销售额(元):
0    1000
1    2500
2    3700
3    5500
4    7500
5    8400
6   10000
dtype: int64

分析:
- 第1天的累计销售就是当天的销售额
- 第2天的累计销售额 = 第1天 + 第2天的销售额
- 以此类推...

Code explanation:

  • Cumulative sum at position 1 = 1000
  • Cumulative sum at position 2 = 1000 + 1500 = 2500
  • Cumulative sum at position 3 = 1000 + 1500 + 1200 = 3700
  • And so on...

Example 2: Handling Data with Missing Values

skipnaThe parameter determines how missing values are handled.

Example

import pandas as pd
import numpy as np

# Create a Series containing missing values
data_with_nan = pd.Series([10, 20, np.nan, 30, 40])

print("Original data (with missing values):")
print(data_with_nan)
print()

# Default skipna=True, skip NaN when calculating cumulative sum
cumulative_skipna = data_with_nan.cumsum()
print("Cumulative sum with skipna=True (default):")
print(cumulative_skipna)
print()

# Set skipna=False, NaN positions will break the cumulative sum
cumulative_no_skipna = data_with_nan.cumsum(skipna=False)
print("Cumulative sum with skipna=False:")
print(cumulative_no_skipna)

Output:

原始数据(含缺失值):
0    10.0
1    20.0
2       NaN
3    30.0
4    40.0
dtype: float64

skipna=True(默认)的累计和:
0    10.0
1    30.0
2       NaN
3    60.0
4   100.0
dtype: float64

skipna=False 的累计和:
0    10.0
1    30.0
2       NaN
3     NaN
4     DataFrame

Code explanation:

  • Whenskipna=TrueWhen set to True, NaN positions are skipped, and subsequent cumulative sums continue based on valid values.
  • Calculation: 10 → 30 → (skip) → 60 → 100
  • Whenskipna=FalseWhen set to False, all positions after encountering NaN become NaN.

Example 3: Analyzing Together with Original Data

In practical analysis, it is often necessary to view both the original data and the cumulative data at the same time.

Example

import pandas as pd

# Create stock price data
stock_prices = pd.Series([100, 102, 98, 105, 103, 108, 110])

# Calculate daily changes
daily_change = stock_prices.diff()

# Calculate cumulative changes
cumulative_change = stock_prices.cumsum() - stock_prices.iloc[0] * len(stock_prices) + stock_prices

# Or more simply, use cumsum to calculate the cumulative change relative to the initial price
print("Stock price data:")
print(stock_prices.values)
print()

# Calculate the price change for the day (compared with the previous day)
print("Daily changes:")
print(daily_change.fillna(0).values)
print()

# Another way: calculate cumulative returns
cumulative_return = ((stock_prices / stock_prices.iloc[0]) - 1) * 100

print("Cumulative returns (%):")
print(cumulative_return.values)
print()

# Cumulative value starting from 100 yuan
initial_price = 100
cumulative_value = initial_price + stock_prices.cumsum() - stock_prices.iloc[0]

print("Cumulative value (assuming initial 100 yuan):")
print(cumulative_value.values)

Output:

股价数据:[100, 102, 98, 105, 103, 108, 110]

每日涨跌(与前一天相比):
[ 0.,  2., -4.,  7., -2.,  5.,  2.]

累计收益(%):
[0.0, 2.0, -2.0, 5.0, 3.0, 8.0, 10.0]

累计价值(假设初始 100 元):
0      100
1      202
2      300
3      405
4      508
5      616
6      726
dtype: int64

分析:第 7 天时,股价从 100 涨到 110,累计收益率为 10%。

Example 4: Practical Application - Monthly Cumulative Data

Demonstrates the typical application of cumsum in financial analysis.

Example

import pandas as pd

# Create monthly sales data
monthly_sales = pd.Series({
    'January': 50000,
    'February': 62000,
    'March': 58000,
    'April': 70000,
    'May': 75000,
    'June': 80000
})

print("Monthly sales (yuan):")
print(monthly_sales)
print()

# Calculate monthly cumulative sales
cumulative_sales = monthly_sales.cumsum()

# Create summary table
summary = pd.DataFrame({
    'Monthly sales': monthly_sales,
    'Cumulative sales': cumulative_sales,
    'Annual target completion ratio': (cumulative_sales / 500000 * 100).round(1)
})

print("Sales summary table:")
print(summary)
print()

print(f"First half cumulative sales: {cumulative_sales.iloc[-1]} yuan")
print(f"Annual target completion rate: {cumulative_sales.iloc[-1]/500000*100:.1f}%")

Output:

月度销售额(元):
1月    50000
2月    62000
累计    112000
4月    70000
5月    75000
6月    80000
dtype: int64

销售汇总表:
              月度销售额   累计销售额  完成年度目标(%)
1月    50000    50000          10.0
2月    62000   112000          22.4
3月    58000   170000          34.0
4月    70000   240000          累计  34.0%
5月    75000   315000          63.0
6月   累计  395000          79.0
dtype: int64

分析:6 月底累计销售额为 39.5 万元,已完成年度目标的 79.0%。

Example 5: Application in DataFrame

cumsum can also be used in a DataFrame column-wise or row-wise.

Example

import pandas as pd

# Create a DataFrame containing sales of multiple products
sales_data = pd.DataFrame({
    'Product A': [100, 150, 120, 180],
    'Product B': [80, 90, 110, 130],
    'Product C': [50, 60, 70, 80]
}, index=['January', 'February', 'March', 'April'])

print("Monthly sales data:")
print(sales_data)
print()

# Calculate cumulative sales by column (default)
cumulative_by_col = sales_data.cumsum()
print("Cumulative sales by product:")
print(cumulative_by_col)
print()

# Calculate cumulative sales by row
cumulative_by_row = sales_data.cumsum(axis=1)
print("Cumulative sales by month:")
print(cumulative_by_row)

Output:

月度销售数据:
    产品A  产品B  产品C
1月  100    80    50
2月  150    90    60
3月 120
4月  180   130    80

按产品累计销量:
    产品A  产品B  产品C
1月  100    80    50
2月  250   170   110
3月 370   280   180
4月 550   410   260

按月度累计销量:
     产品A   产品B   产品C
1月  100.0  180.0  230.0
2月  250.0  340.0  400.0

Code explanation:

  • By default,axis=0the cumulative sum is calculated along the column direction.
  • Settingaxis=1calculates the cumulative sum along the row direction.

Notes

  • cumsum returns a new Series and does not modify the original data.
  • By default, skipna=True, and NaN values are skipped.
  • When skipna=False, encountering NaN will cause all subsequent values to become NaN.
  • In a DataFrame, the axis parameter can be used to control the cumulative direction.
  • The calculation of the cumulative sum is order-dependent and cannot be parallelized.

Summary

Series.cumsum()is a fundamental function in time series analysis. Its main features include:

  • Calculates the cumulative value from the beginning to the current position.
  • Supports handling missing values, controlled via the skipna parameter.
  • Supports row-wise or column-wise calculation in a DataFrame.
  • Widely used in financial analysis, sales statistics, progress tracking, and other scenarios.

The cumulative sum helps us understand the trend of data over time and is an important tool in data analysis. Used together with other cumulative functions (such as cumprod, cummax, cummin), it can provide more comprehensive data insights.

Pandas 常用函数Common Pandas Functions

Other Extensions