Pandas Time Series Analysis
Time series analysis is an important part of data analysis. Pandas provides rich functionality to process and analyze time series data, including resampling, rolling calculations, moving averages, etc.
Basic Time Series Operations
Creating Time Series
Example
import numpy as np
# Create time series data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=100, freq="D")
ts = pd.Series(
np.random.randn(100).cumsum() + 100,
index=dates
)
print("Time series data (first 10 items):")
print(ts.head(10))
print()
# View information
print(f"Index type: {type(ts.index)}")
print(f"Start time: {ts.index.min()}")
print(f"End time: {ts.index.max()}")
print(f"Time span: {ts.index.max() - ts.index.min()}")
Set Date as Index
Example
import numpy as np
# Create DataFrame and set date index
df = pd.DataFrame({
"Date": pd.date_range("2024-01-01", periods=30, freq="D"),
"Sales": np.random.randint(100, 500, 30),
"Visitors": np.random.randint(50, 200, 30)
})
print("Before setting:")
print(df.head())
print()
# Set date as index
df = df.set_index("Date")
print("After setting:")
print(df.head())
print()
# Use loc to query by date
print("Query from 2024-01-05 to 2024-01-10:")
print(df.loc["2024-01-05":"2024-01-10"])
Resampling
Resampling is the process of converting time series data from one frequency to another, including upsampling (increasing data points) and downsampling (decreasing data points).
Downsampling
Example
import numpy as np
# Create daily-level data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=90, freq="D")
ts = pd.Series(np.random.randint(100, 500, 90), index=dates)
print("Daily-level data (first 10 items):")
print(ts.head(10))
print()
# Resample by month (sum)
monthly = ts.resample("M").sum()
print("Monthly sum:")
print(monthly)
print()
# Resample by month (average)
monthly_mean = ts.resample("M").mean()
print("Monthly average:")
print(monthly_mean)
Upsampling
Example
# Create low-frequency data
dates = pd.date_range("2024-01-01", periods=3, freq="M")
ts = pd.Series([100, 200, 150], index=dates)
print("Monthly-level data:")
print(ts)
print()
# Upsample to daily level (needs fill method)
ts_daily = ts.resample("D").ffill()
print("Upsampled to daily level (first 10 items):")
print(ts_daily.head(10))
Common Resampling Methods
| Method | Description |
|---|---|
sum() |
Sum |
mean() |
Mean |
max() / min() |
Max/Min value |
first() / last() |
First/Last value |
count() |
Count of non-null values |
ohlc() |
Open, High, Low, Close |
Rolling Calculations
Rolling calculation is a sliding window calculation on time series data, commonly used to compute moving averages, moving standard deviations, etc.
Moving Average
Example
import numpy as np
# Create data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=30, freq="D")
ts = pd.Series(np.random.randint(100, 200, 30), index=dates)
# Calculate 7-day moving average
rolling_mean = ts.rolling(window=7).mean()
print("7-day moving average (first 10 items):")
print(rolling_mean.head(10))
print()
# Calculate 7-day moving standard deviation
rolling_std = ts.rolling(window=7).std()
print("7-day moving standard deviation:")
print(rolling_std.head(10))
Rolling Apply Custom Function
Example
import numpy as np
ts = pd.Series(range(1, 11))
print("Original data:")
print(ts)
print()
# Rolling sum
print("Rolling 3 sum:")
print(ts.rolling(3).sum())
print()
# Rolling max
print("Rolling 3 max:")
print(ts.rolling(3).max())
print()
# Use apply
print("Rolling 3 custom function (range):")
print(ts.rolling(3).apply(lambda x: x.max() - x.min()))
Time Series Data Visualization
Example
import numpy as np
# Create sample data
np.random.seed(42)
dates = pd.date_range("2024-01-01", periods=100, freq="D")
ts = pd.Series(
np.random.randn(100).cumsum() + 100,
index=dates
)
# Calculate moving average
ma_7 = ts.rolling(7).mean()
ma_30 = ts.rolling(30).mean()
# Display data
print("Time series + moving average:")
print(f"Original data (first 5 items): {ts.head().tolist()}")
print(f"7-day moving average (first 10 items): {ma_7.dropna().head().tolist()}")
print(f"30-day moving average (last 5 items): {ma_30.dropna().tail().tolist()}")
Time Series Feature Extraction
Example
# Create time series
ts = pd.Series(
[1, 2, 3, 4, 5, 6, 7, 8, 9, 10],
index=pd.date_range("2024-01-01", periods=10, freq="D")
)
# Difference
diff = ts.diff()
print("First-order difference:")
print(diff)
print()
# Percentage change
pct = ts.pct_change()
print("Percentage change:")
print(pct)
print()
# Shift
shifted = ts.shift(1)
print("Shift back by 1:")
print(shifted)
Hands-on: Stock Data Analysis
Example
import numpy as np
# Simulate stock data
np.random.seed(42)
n_days = 60
df = pd.DataFrame({
"Date": pd.date_range("2024-01-01", periods=n_days, freq="D"),
"Open price": 100 + np.random.randn(n_days).cumsum(),
"Close price": 100 + np.random.randn(n_days).cumsum(),
"Volume": np.random.randint(1000000, 10000000, n_days)
})
df = df.set_index("Date")
# Calculate daily return
df["Return"] = df["Close price"].pct_change()
# Calculate volatility (7-day rolling standard deviation)
df["Volatility"] = df["Return"].rolling(7).std() * np.sqrt(252) # Annualize
# Calculate moving average line
df["MA5"] = df["Close price"].rolling(5).mean()
df["MA20"] = df["Close price"].rolling(20).mean()
# Generate trading signals (golden cross / death cross)
df["Signal"] = 0
df.loc[df["MA5"] > df["MA20"], "Signal"] = 1
df.loc[df["MA5"] < df["MA20"], "Signal"] = -1
print("Stock data analysis results:")
print(df.tail(10))
Common Issues
1. Dates are not continuous
Some time series do not have data for every day (holidays, etc.), need to useasfreqorreindexto handle.
2. Time zone issues
When handling cross-timezone data, usetz_localizeandtz_convertfor time zone setting and conversion.
3. Impact of missing values
Rolling calculations skip missing values by default, but this may affect the continuity of the results.
Other ExtensionsTime series analysis is a fundamental skill in finance, meteorology, IoT, and other fields. Pandas provides a complete toolchain for handling such data.