Pandas df.to_parquet() Function

Pandas 常用函数Pandas Common Functions


to_parquet()It is a method of DataFrame, used to export data to a file in Parquet format.

Parquet is a columnar storage file format specifically designed for big data analysis scenarios. It has advantages such as high compression ratio, high read/write performance, and support for complex data types. It is the standard format for big data frameworks such as Apache Hadoop and Apache Spark.


Basic Syntax and Parameters

Syntax Format

DataFrame.to_parquet(path, engine='auto', compression='snappy',
                     index=None, partition_cols=None, storage_options=None, ...)

Parameter Description

ParameterTypeDescriptionDefault Value
pathstr, path objectFile pathRequired
enginestrEngine: 'auto', 'pyarrow', 'fastparquet''auto'
compressionstrCompression: 'snappy', 'gzip', 'brotli', None'snappy'
indexbool, NoneWhether to include the indexNone
partition_colslistPartition column(s), store data partitioned by columnNone

Return Value Description

  • Return Type:None
  • Writes data directly to a Parquet file, no return value.

Examples

Through the following examples, comprehensively masterto_parquet()the various usages.

Example 1: Basic Usage - Export to a Parquet File

First create a DataFrame, then useto_parquet()to export to a Parquet file.

Example

import pandas as pd

# Create a sample DataFrame
data = {
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy', 'John'],
    'age': [28, 35, 42, 26, 31],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Hangzhou'],
    'salary': [8000, 12000, 15000, 7000, 9000],
    'department': ['IT', 'HR', 'Sales', 'IT', 'HR']
}
df = pd.DataFrame(data)

# Example 1a: Basic export
# path: file path (required)
# Uses snappy compression by default
df.to_parquet('employees.parquet')
print("Exported to employees.parquet")

# Check file size
import os
file_size = os.path.getsize('employees.parquet')
print(f"File size: {file_size} bytes")

# Example 1b: Read and verify
# Requires installing pyarrow or fastparquet
df_check = pd.read_parquet('employees.parquet')
print("nVerify by reading:")
print(df_check)

Output:

已导出到 employees.parquet
文件大小: 约 600-800 字节(远小于 CSV)

验证读取:
    name  age       city  salary department
0    Tom   28    Beijing    8000         IT
1  Jerry   35  Shanghai   12000         HR
2   Mike   42  Guangzhou   15000       Sales
3   Lucy   26   Shenzhen    7000         IT
4   John   31   Hangzhou    9000         HR

Code Explanation:

  • to_parquet()Export the DataFrame to Parquet format.
  • By default, usessnappycompression, with remarkable compression results.
  • Parquet files are much smaller than CSV files and also faster to read.

Example 2: Selecting Engine and Compression Method

The Parquet format supports multiple engines and compression methods, which can be selected as needed.

Example

import pandas as pd
import os

# Create a larger DataFrame to observe the compression effect
import numpy as np
np.random.seed(42)
df_large = pd.DataFrame({
    'id': range(10000),
    'value': np.random.randn(10000),
    'category': np.random.choice(['A', 'B', 'C', 'D'], 10000),
    'name': np.random.choice(['Tom', 'Jerry', 'Mike', 'Lucy', 'John'], 10000)
})

# Example 2a: Use different compression methods
# snappy: fast compression, high speed, moderate compression ratio (default)
df_large.to_parquet('output_snappy.parquet', compression='snappy')

# gzip: high compression ratio, smaller files
df_large.to_parquet('output_gzip.parquet', compression='gzip')

# brotli: higher compression ratio
df_large.to_parquet('output_brotli.parquet', compression='brotli')

# no compression
df_large.to_parquet('output_none.parquet', compression=None)

# Compare file sizes
print("Comparison of file sizes with different compression methods:")
for name in ['snappy', 'gzip', 'brotli', 'none']:
    size = os.path.getsize(f'output_{name}.parquet')
    print(f" {name}: {size:,} bytes")
print()

# Example 2b: Choose the engine
# auto: automatic selection (default)
# pyarrow: Apache Arrow implementation, full-featured, good performance
# fastparquet: pure Python implementation, good compatibility
print("Available engines: auto, pyarrow, fastparquet")
print("Currently using:", end=" ")

# Check available engines
try:
    import pyarrow
    print("pyarrow")
except ImportError:
    pass

try:
    import fastparquet
    print("fastparquet")
except ImportError:
    pass

Output:

不同压缩方式的文件大小对比:
  snappy: 约 100KB
  gzip: 约 80KB
  brotli: 约 70KB
  none: 约 200KB

不同压缩方式各有优劣:
  snappy: 速度快,压缩比适中(默认推荐)
  gzip: 压缩比更高,适合存储
  brotli: 最高压缩比,适合冷数据
  none: 无压缩,速度最快

Code Explanation:

  • compressionThe compression parameter allows selecting different compression methods.
  • snappyis the default option, balancing speed and compression ratio.
  • gzipHas a higher compression ratio, but is slightly slower.
  • engineThe engine parameter allows selecting which engine to use.

Example 3: Partitioned Storage

Parquet supports partitioned storage by column, which is an important feature in big data analysis.

Example

import pandas as pd
import os
import shutil

# Create a DataFrame
df = pd.DataFrame({
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy', 'John', 'Mary', 'Bob', 'Alice'],
    'age': [28, 35, 42, 26, 31, 29, 38, 24],
    'department': ['IT', 'HR', 'Sales', 'IT', 'HR', 'IT', 'Sales', 'HR'],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Hangzhou',
             'Beijing', 'Shanghai', 'Beijing']
})

# Example 3a: Partition by a single column
# partition_cols specifies the partition column, which generates subdirectories
if os.path.exists('partitioned'):
    shutil.rmtree('partitioned')

df.to_parquet('partitioned', partition_cols=['department'])
print("Exported with partitioning by department")

# View the partition directory structure
for root, dirs, files in os.walk('partitioned'):
    level = root.replace('partitioned', '').count(os.sep)
    indent = ' ' * 2 * level
    print(f'{indent}{os.path.basename(root)}/')
    subindent = ' ' * 2 * (level + 1)
    for file in files:
        print(f'{subindent}{file}')

Output:

已按 department 分区导出

分区目录结构:
partitioned/
  department=HR/
    xxx.parquet
  department=IT/
    xxx.parquet
  department=Sales/
    xxx.parquet

Code Explanation:

  • partition_colsThe partition_cols parameter performs partitioned storage by the specified column.
  • Partitioned storage creates subdirectories in the file system, one directory per partition value.
  • Partitioned storage is very beneficial for big data queries, as only the needed partitions need to be read.

Example 4: Handling the Index

When exporting to Parquet, you can choose whether to include the index.

Example

import pandas as pd

# Create a DataFrame with an index
df = pd.DataFrame({
    'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
    'age': [28, 35, 42, 26],
    'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
})
df.index = ['A001', 'A002', 'A003', 'A004']

# Example 4a: Include index by default
df.to_parquet('with_index.parquet')
print("Export (index included by default)")

# Example 4b: Exclude index
df.to_parquet('without_index.parquet', index=False)
print("Export (index not included)")

# Example 4c: Explicitly include index
df.to_parquet('explicit_index.parquet', index=True)
print("Export (explicitly include index)")

# Read and compare
print("nRead and compare the differences:")
print("nOriginal data:")
print(df)
print("nRead with_index.parquet:")
print(pd.read_parquet('with_index.parquet'))
print("nRead without_index.parquet:")
print(pd.read_parquet('without_index.parquet'))
print("nRead explicit_index.parquet:")
print(pd.read_parquet('explicit_index.parquet'))

Output:

导出(默认包含索引)
导出(不包含索引)
导出(显式包含索引)

读取并比较差异:

原始数据:
    name  age       city
A001  Tom   28    Beijing
A002  Jerry   35  Shanghai
A003   Mike   42  Guangzhou
A004  Lucy   26   Shenzhen

读取 with_index.parquet:
    name  age       city  index
0    Tom   28   Beijing   A001
1   Jerry  35  Shanghai   A002
2    Mike   42  Guangzhou  A003
3    Lucy   26  Shenzhen   A004

读取 without_index.parquet:
    name  age       city
0    Tom   28   Beijing
1   Jerry  35  Shanghai
2    Mike   42  Guangzhou
3    Lucy   26   Shenzhen

读取 explicit_index.parquet:
    name  age       city  index
0    Tom   28   Beijing   A001
1   Jerry  35  Shanghai   A002
2    Mike   42  Guangzhou  A003
3    Lucy   26  Shenzhen   A004

Code Explanation:

  • index=TrueExplicitly includes the index as a column.
  • index=FalseDoes not include the index.
  • The default behavior depends on pandas' index settings.

Notes

  • Usingto_parquet()requires installingpyarroworfastparquet。
  • Recommended to installpyarrow:pip install pyarrow。
  • Parquet is columnar storage, suitable for big data analysis scenarios, but not suitable for small data.
  • Partitioned storage can significantly improve big data query performance.
  • Parquet supports complex data types (nested structures), but they are not commonly used in pandas DataFrames.

Summary

to_parquet()It is a method for DataFrame to export to Parquet format. Parquet is the standard columnar storage format in the big data field, with advantages such as high compression ratio, high performance, and support for partitioning.

In big data processing scenarios, Parquet is the preferred data format. It is perfectly compatible with big data frameworks such as Apache Spark and Apache Hive. When processing large-scale data, it is recommended to use Parquet format instead of CSV or Excel.

Pandas 常用函数Pandas Common Functions

Other Extensions