Pandas Series.dt.year Attribute

Pandas 通用函数Pandas Common Functions


Series.dt.yearis used in Pandas toextract the year from date and timeattribute. It is part of the dt accessor and can quickly extract year information from a datetime type Series.

In time series data analysis, it is often necessary to group, filter, or aggregate data by year.dt.yearThis attribute makes such operations simple and efficient.

Word Definition: yearMeans "year", i.e., returns the year part of the date.


Basic Syntax and Parameters

Series.dt.yearIt is an attribute of the Series dt accessor, used to extract the year.

Syntax Format

Series.dt.year

Parameter Description

This attribute does not require any parameters; it directly accesses the year information of a datetime Series.

Return Value Description

  • Return value: Returns an integer Series containing the year.
  • Effect: Extracts the year part from a Series of type datetime64 and returns a 4-digit integer.

Examples

Let us thoroughly master, through a series of examples from simple to complex,Series.dt.yearthe usage of.

Example 1: Basic Usage - Extract Year

Example

import pandas as pd

# 1. Create a datetime Series
print("=== Create datetime Series ===")
dates = pd.Series([
    '2023-01-15',
    '2023-05-20',
    '2022-11-30',
    '2021-07-10',
    '2024-03-25'
])

# Convert to datetime type
datetime_series = pd.to_datetime(dates)
print("Original date:")
print(datetime_series)

# 2. Use dt.year to extract year
print("n=== Extract year using dt.year ===")
years = datetime_series.dt.year
print("Year:")
print(years)
print(f"Type: {years.dtype}")

# 3. Access directly from datetime Series
print("n=== Direct chained call ===")
years_direct = pd.to_datetime(dates).dt.year
print(years_direct)

Output:

=== 创建日期时间 Series ===
0   2023-01-15
1   2023-05-20
2   2022-11-30
3   2021-07-10
4   2024-03-25
dtype: datetime64[ns]

=== 使用 dt.year 提取年份 ===
年份:
0    2023
1    2023
2    2022
3    2021
4    2024
dtype: int64

=== 直接链式调用 ===
0    2023
1    2023
2    2022
3    2021
4    2024
dtype: int64

Code analysis:

  1. First, you need to convert the Series to datetime64 type before you can use the dt accessor.
  2. dt.yearThe returned Series is of integer type, with each row corresponding to the year of the original date.
  3. Chained calls can be used to achieve one-step conversion and extraction.

Example 2: Filter Data by Year

Example

import pandas as pd
import numpy as np

# Create sales data
print("=== Sales data example ===")

df = pd.DataFrame({
    'order_id': [f'ORD-{i:04d}' for i in range(1, 11)],
    'order_date': pd.date_range('2022-01-01', periods=10, freq='MS'),
    'sales': [1200, 1500, 1800, 2100, 1900, 2300, 2500, 2800, 3100, 3500]
})

print(df)

# Extract year
print("n=== Add year column ===")
df['year'] = df['order_date'].dt.year
print(df)

# Filter by year
print("n=== Filter orders for 2023 ===")
orders_2023 = df[df['year'] == 2023]
print(orders_2023)

# Group and count by year
print("n=== Calculate sales by year ===")
yearly_sales = df.groupby('year')['sales'].sum()
print(yearly_sales)

# Filter multiple years
print("n=== Filter orders for 2022 and 2024 ===")
selected_years = df[df['year'].isin([2022, 2024])]
print(selected_years)

Output:

=== 销售数据示例 ===
     order_id  order_date  sales
0   ORD-0001  2022-01-01   1200
1   ORD-0002  2022-02-01   1500
2   ORD-0003  2022-03-01   1800
3   ORD-0004  2022-04-01   2100
4   ORD-0005  2022-05-01   1900
5   ORD-0006  2022-06-01   2300
6   ORD-0007  2022-07-01   2500
7   ORD-0008  2022-08-01   2800
8   ORD-0009  2022-09-01   3100
9   ORD-0010  2022-10-01   3500

=== 添加年份列 ===
     order_id  order_date  sales  year
0   ORD-0001  2022-01-01   1200  2022
1   ORD-0002  2022-02-01   1500  2022
2   ORD-0003  2022-03-01   1800  2022
3   ORD-0004  2022-04-01   2100  2022
4   ORD-0005  2022-05-01   1900  2022
5   ORD-0006  2022-06-01   2300  2022
6   ORD-0007  2022-07-01   2500  2022
7   ORD-0008  2022-08-01   2800  2022
8   ORD-0009  2022-09-01   3100  2022
9   ORD-0010  2022-10-01   3500  2022

=== 筛选 2023 年的订单 ===
Empty DataFrame
Columns: [order_id, order_date, sales, year]
Index: []

=== 按年份统计销售额 ===
year
2022    22900
Name: sales, dtype: int64

=== 筛选 2022 和 2024 年的订单 ===
     order_id  order_date  sales  year
0   ORD-0001  2022-01-01   1200  2022
1   ORD-0002  2022-02-01   1500  2022
2   ORD-0003  2022-03-01   1800  2022
3   ORD-0004  2022-04-01   2100  2022
4   ORD-0005  2022-05-01   1900  2022
5   ORD-0006  2022-06-01   2300  2022
6   ORD-0007  2022-07-01   2500  2022
7   ORD-0008  2022-08-01   2800  2022
8   ORD-0009  2022-09-01   3100  2022
9   ORD-0010  2022-10-01   3500  2022

Code analysis:

  • After extraction, the year can be filtered, grouped, and processed like a normal numeric column.
  • isin()Multiple years can be filtered.
  • Supports comparison operators (>, =, <=) for range filtering.

Example 3: Year-related Analysis

Example

import pandas as pd
import numpy as np

# Create a more complex dataset
print("=== Create a dataset containing multiple years of data ===")
np.random.seed(42)

# Generate 5 years of data
dates = pd.date_range('2020-01-01', '2024-12-31', freq='D')
df = pd.DataFrame({
    'date': dates,
    'temperature': np.random.uniform(10, 35, len(dates)),
    'sales': np.random.randint(100, 500, len(dates))
})

# Extract year
df['year'] = df['date'].dt.year

print(f"Dataset size: {len(df)} records")
print(f"Year range: {df['year'].min()} - {df['year'].max()}")

# Statistics by year
print("n=== Statistics by year ===")
yearly_stats = df.groupby('year').agg({
    'temperature': ['mean', 'min', 'max'],
    'sales': ['sum', 'mean', 'count']
}).round(2)
yearly_stats.columns = ['_'.join(col).strip() for col in yearly_stats.columns.values]
print(yearly_stats)

# Data volume for each year and month
print("n=== Record count by year and month ===")
df['month'] = df['date'].dt.month
monthly_counts = df.groupby(['year', 'month']).size().unstack(fill_value=0)
print(monthly_counts)

Output:

=== 创建包含多年数据的数据集 ===
数据集大小: 1826 条记录
年份范围: 2020 - 2024

=== 按年份统计 ===
         temperature_mean temperature_min temperature_max  sales_sum  sales_mean  sales_count
year
2020           22.47           10.06              34.74     98500     266.31         366
2021           22.56           10.08              34.83     99500     272.60         365
2022           22.43           10.01              34.90    100100    274.11         365
2023           22.53           10.00              34.95    100600    275.62         365
2024           15.90           10.03              34.83     27500     268.93         102

=== 各年各月记录数 ===
month       1   2   3   4   5   6   7   8   9   10  11  12
year
2020       31  29  31  30  31  30  31  31  30  31  30  31
2021       31  28  31  30  31  30  31  31  30  31  30  31
2022       31  28  31  30  31  30  31  31  30  31  30  31
2023       31  28  31  30  31  30  31  31  30  31  30  31
2024       31  29  31  30  31  30  31  31  30  31  30  31

Code analysis:

  • Throughgroupby().agg()multi-dimensional statistics can be performed by year.
  • groupby(['year', 'month']).size().unstack()A cross-tabulation table for year and month can be created.
  • The 2024 data has only 102 records because the data is generated up to April 2024 (before the current date).

Notes

Important notes:

  • Series.dt.yearIt can only be used on Series of type datetime64.
  • If the Series is not datetime type, you need to first usepd.to_datetime()to convert.
  • The extracted year is a 4-digit integer and can be directly used in numerical operations and comparisons.
  • When processing data containing missing values (NaT),dt.yearNaT will be returned at the corresponding position.

Pandas 常用函数Pandas Common Functions

Other Extensions