Pandas pd.date_range() Function

Pandas 通用函数Pandas Common Functions


pd.date_range()is used in the Pandas library togenerate date rangesfunction. It creates a DatetimeIndex containing consecutive dates, commonly used for creating, indexing, and aligning time series data.

In time series analysis, it is often necessary to generate date sequences with a fixed frequency, such as daily, monthly, or yearly data. pd.date_range() provides flexible parameters to meet various frequency requirements.

Term Definition: date_rangeIt means "date range", i.e., generating a collection of consecutive dates.


Basic Syntax and Parameters

pd.date_range()is a top-level function of the Pandas library, used to generate datetime indices with a specified frequency.

Syntax Format

pd.date_range(start=None, end=None, periods=None, freq=None, tz=None, normalize=False, name=None, closed=None)

Parameter Description

Parameter Type Required Description Default Value
start String or datetime Optional Start date. None
end String or datetime Optional End date. None
periods Integer Optional Number of dates to generate. Use any two of start, end, and periods. None
freq String or DateOffset Optional Frequency: 'D' (day), 'H' (hour), 'M' (month end), 'MS' (month start), 'Y' (year end), 'W' (week), etc. 'D'
tz String Optional Time zone name, e.g., 'UTC', 'Asia/Shanghai'. None
normalize Boolean Optional If True, normalizes times to midnight. False
name String Optional Name of the resulting DatetimeIndex. None
closed String Optional Interval inclusivity: 'left', 'right', None (default, both inclusive). None

Return Value Description

  • Return value: Returns a DatetimeIndex object containing consecutive date-times.
  • Effect: Generates a date sequence based on the specified start date, end date, and frequency.

Examples

Let us thoroughly master, through a series of examples from simple to complex,pd.date_range()its usage.

Example 1: Basic Usage - Generating Date Range

Example

import pandas as pd

# 1. Use the start and end parameters to generate a date range (default frequency is daily)
print("=== From 2023-01-01 to 2023-01-10 ===")
dates = pd.date_range(start='2023-01-01', end='2023-01-10')
print(dates)
print(f"Type: {type(dates)}")
print(f"Data type: {dates.dtype}")

# 2. Use the periods parameter to specify the number to generate
print("n=== Generate 7 days of dates ===")
week_dates = pd.date_range(start='2023-01-01', periods=7)
print(week_dates)

# 3. Use the freq parameter to specify the frequency
print("n=== One date per week (freq='W') ===")
weekly_dates = pd.date_range(start='2023-01-01', periods=5, freq='W')
print(weekly_dates)

# 4. Date at the beginning of each month (freq='MS' = Month Start)
print("n=== Month start (freq='MS') ===")
monthly_dates = pd.date_range(start='2023-01-01', periods=6, freq='MS')
print(monthly_dates)

# 5. Date at the end of each year (freq='YE' = Year End)
print("n=== Year end (freq='YE') ===")
yearly_dates = pd.date_range(start='2023-01-01', periods=3, freq='YE')
print(yearly_dates)

Output:

=== 从 2023-01-01 到 2023-01-10 ===
DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-05',
               '2023-01-06', '2023-01-07', '2023-01-08', '2023-01-09', '2023-01-10'],
              dtype='datetime64[ns]', freq='D')
Type: <class 'pandas.core.indexes.datetimes.DatetimeIndex'>
数据类型: datetime64[ns]

=== 生成 7 天的日期 ===
DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-05',
               '2023-01-06', '2023-01-07'], dtype='datetime64[ns]', freq='D')

=== 每周一个日期 (freq='W') ===
DatetimeIndex(['2023-01-01', '2023-01-08', '2023-01-15', '2023-01-22', '2023-01-29'], dtype='datetime64[ns]', freq='W-SUN')

=== 每月初 (freq='MS') ===
DatetimeIndex(['2023-01-01', '2023-02-01', '2023-03-01', '2023-04-01', '2023-05-01',
               '2023-06-01'], dtype='datetime64[ns]', freq='MS')

=== 每年末 (freq='YE') ===
DatetimeIndex(['2023-12-31', '2024-12-31', '2025-12-31'], dtype='datetime64[ns]', freq='YE-DEC')

Code Explanation:

  1. pd.date_range()Generates a date sequence with daily frequency by default.
  2. freqThe parameter supports multiple frequency abbreviations: 'D' (day), 'W' (week), 'M' (month end), 'MS' (month start), 'Y' (year end), 'H' (hour), etc.
  3. What is returned is a DatetimeIndex object, which can be directly used as the index of a Series or DataFrame.

Example 2: Generating Time Series Data

Create a DataFrame with a date index for time series analysis.

Example

import pandas as pd
import numpy as np

# 1. Create a DataFrame with a date index
print("=== Create a simple time series ===")

# Generate 30 days of dates
date_index = pd.date_range(start='2023-01-01', periods=30, freq='D')

# Create DataFrame
df = pd.DataFrame({
    'date': date_index,
    'value': np.random.randint(10, 100, size=30),  # Random values
    'category': ['A', 'B'] * 15  # Alternating categories
})

print(df.head(10))

# 2. Set date as index
print("n=== Set date as index ===")
df_indexed = df.set_index('date')
print(df_indexed.head())

# 3. Generate hourly data (for log analysis)
print("n=== Hourly time series ===")
hourly_index = pd.date_range(start='2023-01-01 08:00', periods=24, freq='H')
hourly_data = pd.DataFrame({
    'timestamp': hourly_index,
    'visitors': np.random.randint(0, 100, size=24)
})
print(hourly_data)

Output:

=== 创建一个简单的时间序列 ===
        date  value category
0 2023-01-01     67        A
1 2023-01-02     23        B
2 2023-01-03     45        A
3 2023-01-04     89        B
4 2023-01-05     12        A
5 2023-01-06     34        B
6 2023-01-07     56        A
7 2023-01-08     78        B
8 2023-01-09     91        A
9 2023-01-10     55        B

=== 设置日期为索引 ===
             value category
date
2023-01-01     67        A
2023-01-02     23        B
2023-01-03     45        A
2023-01-04     89        B
2023-01-05     12        A
2023-01-06     34        B

=== 每小时的时间序列 ===
                 timestamp  visitors
0  2023-01-01 08:00:00      25
1  2023-01-01 09:00:00      41
2  2023-01-01 10:00:00      77
3  2023-01-01 11:00:00      12
4  2023-01-01 12:00:00      58
5  2023-01-01 13:00:00      33
6  2023-01-01 14:00:00      45
7  2023-01-01 15:00:00      67
8  2023-01-01 16:00:00      29
9  2023-01-01 17:00:00      51
10 2023-01-01 18:00:00      44
11 2023-01-01 19:00:00      36
12 2023-01-01 20:00:00      28
13 2023-01-01 21:00:00      62
14 2023-01-01 22:00:00      19
15 2023-01-01 23:00:00      73
16 2023-01-02 00:00:00      41
17 2023-01-02 01:00:00      55
18 2023-01-02 02:00:00      38
19 2023-01-02 03:00:00      27
20 2023-01-02 04:00:00      49
21 2023-01-02 05:00:00      63
22 2023-01-02 06:00:00      34
23 2023-01-02 07:00:00      56

Code Explanation:

  • A DatetimeIndex can be directly used as the index of a DataFrame.
  • DataFrames with a date index support rich time series operations, such as resampling and rolling calculations.
  • freq='H'Hourly-level data can be generated, suitable for scenarios such as log analysis.

Example 3: Time Zone and closed Parameter

Example

import pandas as pd

# 1. Use the tz parameter to specify the time zone
print("=== Date range with time zone ===")
dates_tz = pd.date_range(start='2023-01-01', periods=3, freq='D', tz='Asia/Shanghai')
print(dates_tz)
print(f"Time zone: {dates_tz.tz}")

# 2. Use UTC time zone
print("n=== UTC time zone ===")
dates_utc = pd.date_range(start='2023-01-01', periods=3, freq='D', tz='UTC')
print(dates_utc)

# 3. Use the closed parameter to control interval inclusivity
print("n=== Comparison of closed parameter ===")
print("closed=None (default, both ends inclusive):", pd.date_range('2023-01-01', '2023-01-03', freq='D', closed=None))
print("closed='left':", pd.date_range('2023-01-01', '2023-01-03', freq='D', closed='left'))
print("closed='right':", pd.date_range('2023-01-01', '2023-01-03', freq='D', closed='right'))

# 4. Use the normalize parameter to normalize time
print("n=== normalize parameter ===")
dates_with_time = pd.date_range(start='2023-01-01 14:30:00', periods=3, freq='H')
print("Before normalization:", dates_with_time)

dates_normalized = pd.date_range(start='2023-01-01 14:30:00', periods=3, freq='H', normalize=True)
print("After normalization:", dates_normalized)

Output:

=== 带时区的日期范围 ===
DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03'], dtype='datetime64[ns, Asia/Shanghai]')
时区: Asia/Shanghai

=== UTC 时区 ===
DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03'], dtype='datetime64[ns, UTC]')

=== closed 参数对比 ===
closed=None (默认, 两端都闭): DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03'])
closed='left': DatetimeIndex(['2023-01-01', '2023-01-02'])
closed='right': DatetimeIndex(['2023-01-02', '2023-01-03'])

=== normalize 参数 ===
归一化前: DatetimeIndex(['2023-01-01 14:30:00', '2023-01-01 15:30:00', '2023-01-01 16:30:00'])
归一化后: DatetimeIndex(['2023-01-01', '2023-01-02', '2023-01-03'])

Code Explanation:

  • tzThe parameter can specify the time zone, which is very useful when handling data across time zones.
  • closedThe parameter can control the inclusivity of the date interval, with both ends inclusive by default.
  • normalize=TrueNormalizes times to midnight, making it convenient to aggregate data by date.

Notes

Important notes:

  • You must specifystartandend, or specifystartandperiods, or specifyendandperiods, but you cannot specify only one at the same time.
  • Frequency abbreviations are case-sensitive; for example, 'D' means day, and 'd' will raise an error.
  • When handling time zone data, ensure all dates use a uniform time zone; otherwise, unexpected results may occur.
  • When generating a large number of dates (e.g., multi-year data), pay attention to memory usage.

Pandas 通用函数Pandas Common Functions

Other Extensions