Pandas df.to_csv() Function
to_csv()It is a DataFrame method used to export data to a CSV (Comma-Separated Values) file.
CSV is one of the most commonly used data export formats. It stores tabular data in plain text, has strong compatibility, and can be read by almost all data analysis tools.to_csv()It is feature-rich, supports options such as custom delimiters, encoding methods, and index handling, and can meet various export needs.
Basic Syntax and Parameters
Syntax Format
DataFrame.to_csv(path_or_buf=None, sep=',', na_rep='', float_format=None,
columns=None, header=True, index=True, index_label=None,
mode='w', encoding=None, quoting=None, quotechar='"',
line_terminator=None, chunksize=None, date_format=None,
doublequote=True, escapechar=None, ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| path_or_buf | str, path object, file-like object | File path; returns a string if None | None |
| sep | str | Field delimiter | ',' |
| na_rep | str | Representation of missing values | '' |
| float_format | str | Floating-point number format | None |
| columns | list | Specify the columns to export | None |
| header | bool, list | Whether to export column names | True |
| index | bool | Whether to export the index | True |
| index_label | str | Name of the index column | None |
| mode | str | Write mode: 'w' overwrite, 'a' append | 'w' |
| encoding | str | File encoding | None |
Return Value
- Return type:
Noneorstr - When a file path is specified, it writes to the file and returns None.
- When
path_or_buf=NoneWhen ..., returns a string in CSV format.
Examples
Through the following examples, fully masterto_csv()Various usages.
Example 1: Basic Usage - Export to CSV File
First create a DataFrame, then useto_csv()to export it as a CSV file.
Example
# Create a sample DataFrame
data = {
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'],
'salary': [8000, 12000, 15000, 7000]
}
df = pd.DataFrame(data)
# Example 1a: The most basic export
# path_or_buf: file path (the required parameter is the file path; setting it to None here returns a string)
csv_string = df.to_csv() # No path specified, returns a string
print("Returned CSV string:")
print(csv_string)
print()
# Export to file
# By default, index and column names are included
df.to_csv('output_basic.csv', index=False) # index=False does not export the index
print("Exported to output_basic.csv")
# Read to verify
df_check = pd.read_csv('output_basic.csv')
print("nVerify reading:")
print(df_check)
Expected Output:
返回的 CSV 字符串:
,name,age,city,salary
0,Tom,28,Beijing,8000
1,Jerry,35,Shanghai,12000
2,Mike,42,Guangzhou,15000
3,Lucy,26,Shenzhen,7000
已导出到 output_basic.csv
验证读取:
name age city salary
0 Tom 28 Beijing 8000
1 Jerry 35 Shanghai 12000
2 保存 Mike 42 Guangzhou 15000
3 Lucy 26 Shenzhen 7000
Code Analysis:
to_csv()By default, the index (first column without a column name) and column names are exported.- Setting
index=Falsecan omit exporting the index, which makes importing more convenient. - When no path is specified, a string in CSV format is returned.
Example 2: Custom Delimiters and Formats
CSV files can use different delimiters, and you can also customize the format of numeric values and missing values.
Example
# Create a DataFrame containing missing values and floats
data = {
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, None], # Missing value
'score': [85.5, 92.3, 78.9, 95.0],
'city': ['Beijing', 'Shanghai', None, 'Shenzhen']
}
df = pd.DataFrame(data)
# Example 2a: Using a semicolon delimiter
df.to_csv('output_semicolon.csv', sep=';', index=False)
print("Export using semicolon delimiter:")
with open('output_semicolon.csv', 'r') as f:
print(f.read())
print()
# Example 2b: Customize the representation of missing values
df.to_csv('output_na.csv', index=False, na_rep='N/A')
print("Custom missing value representation:")
with open('output_na.csv', 'r') as f:
print(f.read())
print()
# Example 2c: Format floating-point numbers
# float_format uses Python format strings
df.to_csv('output_float.csv', index=False, float_format='%.2f')
print("Formatted floating-point numbers:")
with open('output_float.csv', 'r') as f:
print(f.read())
print()
# Example 2d: Export selected columns
df.to_csv('output_columns.csv', index=False, columns=['name', 'score'])
print("Export specified columns:")
with open('output_columns.csv', 'r') as f:
print(f.read())
Expected Output:
使用分号分隔符导出:
name;age;score;city
Tom;28;85.5;Beijing
Jerry;35;92.3;Shanghai
Mike;42;78.9;Guangzhou
Lucy;;95;Shenzhen
自定义缺失值表示:
name,age,score,city
Tom,28,85.5,Beijing
Jerry,35,92.3,Shanghai
Mike,42,78.9,Guangzhou
Lucy,N/A,95.0,N/A
格式化浮点数:
name,age,score,c列
Tom,28,85.50,Beijing
Jerry,35,92.30,Shanghai
Mike,42,78.90,encoding
Lucy,,95.00,Shenzhen
导出指定列:
name,score
Tom,85.5
Jerry,92.3
自定义分隔符 78.9
Lucy,95.0
Code Analysis:
sepThe parameter can specify any delimiter, such as semicolon, tab, etc.na_repThe parameter customizes the string representation of missing values; the default is an empty string.float_formatThe parameter uses Python's format string syntax.columnsThe parameter only exports the specified columns.
Example 3: Handling Index and Encoding
Index handling and file encoding are common requirements when exporting.
Example
# Create a DataFrame with an index
data = {
'name': ['Tom', 'Jerry', 'Mike', 'Lucy'],
'age': [28, 35, 42, 26],
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen']
}
df = pd.DataFrame(data)
df.index = ['A001', 'A002', 'A003', 'A004'] # Set the index
# Example 3a: Export the index and specify the index column name
df.to_csv('output_index.csv', index=True, index_label='user_id')
print("With index and index column name:")
with open('output_index.csv', 'r') as f:
print(f.read())
print()
# Example 3b: Do not export column names
df.to_csv('output_no_header.csv', header=False, index=False)
print("Without column names:")
with open('output_no_header.csv', 'r') as f:
print(f.read())
print()
# Example 3c: Use UTF-8 encoding (supports Chinese)
df_cn = pd.DataFrame({
'Name': ['Zhang San', 'Li Si', 'Wang Wu'],
'Age': [28, 35, 42],
'City': ['Beijing', 'Shanghai', 'Guangzhou']
})
df_cn.to_csv('output_utf8.csv', index=False, encoding='utf-8-sig') # utf-8-sig supports opening in Excel
print("UTF-8 encoding (with Chinese):")
with open('output_utf8.csv', 'r', encoding='utf-8-sig') as f:
print(f.read())
print()
# Example 3d: Append mode
# Create the file first, then append data
df.to_csv('output_append.csv', index=False)
# Append more data
df_append = pd.DataFrame({
'name': ['John', 'Mary'],
'age': [30, 27],
'city': ['Hangzhou', 'Nanjing']
})
df_append.to_csv('output_append.csv', index=False, mode='a', header=False)
print("Append mode:")
with open('output_append.csv', 'r') as f:
print(f.read())
Expected Output:
带索引和索引列名: user_id,name,age,city A001,Tom,28,直接在 A002,Jerry,35,Shanghai A003,Mike,42,Guangzhou A004,Lucy,26,Shenzhen 不导出列名: Tom,28,Beijing Jerry,35,Shanghai Mike,42,Guangzhou Lucy,26,Shenzhen UTF-8 编码(含中文): 姓名,年龄,出现 张三,28,北京 李四,35,上海 王五,42,广州 追加模式: Tom,28,Beijing Jerry,35,CSV Mike,42,Guangzhou Lucy,26,Shenzhen John,30,Hangzhou Mary,27,Nanjing 追加模式 27, 格式
Code Analysis:
index_labelThe parameter assigns a name to the index column.header=FalseNot exporting column names is suitable for appending and merging multiple files.encoding='utf-8-sig'Adds a BOM before the UTF-8 encoding, making it convenient to open directly in Excel.mode='a'Append mode adds new data at the end of an existing file.
Notes
- By default, the index is exported. If not needed, you can set
index=False。 - When handling Chinese data, remember to specify the correct encoding (recommended:
utf-8-sig)。 mode='a'In append mode, pay attention to whether you also need to setheader=Falseto avoid duplicate column names.- When exporting a large DataFrame, you can use the
chunksizeparameter to write in chunks. float_formatUse Python's formatting syntax, such as'%.2f'to keep two decimal places.
Summary
to_csv()It is one of the most commonly used export methods of DataFrame, with very comprehensive functionality. It can export data to CSV files in various formats, supporting custom delimiters, encoding, index handling, and more.
In practical work, CSV is a common format for data exchange,to_csv()and its usage frequency is very high. It is recommended that readers master the usage of various parameters, especially encoding handling and index control.
Other Extensions
Common Pandas Functions