Pandas pd.ead_csv() Function
read_csv()is the most commonly used data reading function in the pandas library, used to read data from CSV (Comma-Separated Values) files and create a DataFrame.
CSV files are a simple and widely used data exchange format that stores tabular data in plain text, where each row represents a record and fields are separated by commas.read_csv()It can intelligently parse CSV files, automatically recognize column names and data types, and handle various delimiter and encoding issues.
Basic Syntax and Parameters
Syntax Format
pandas.read_csv(filepath_or_buffer, sep=',', delimiter=None, header='infer',
names=None, index_col=None, usecols=None, dtype=None,
skiprows=None, nrows=None, na_values=None, ...)
Parameter Description
| Parameter | Type | Description | Default Value |
|---|---|---|---|
| filepath_or_buffer | str, path object, or file-like object | The path, URL, or file object of the CSV file | Required |
| sep | str | Field delimiter, comma by default for CSV | ',' |
| header | int, list of int, 'infer' | Row number used as column names, 0 means the first row | 'infer' |
| names | list-like | Custom list of column names | None |
| index_col | int, str, list of int, list of str, False | Column used as the row index | None |
| usecols | list-like, callable | Only read the specified columns | None |
| dtype | dict | Data types for specified columns, e.g., {'a': np.float64} | None |
| skiprows | list-like, int | Skip the specified rows | None |
| nrows | int | Only read the first n rows | None |
| na_values | scalar, str, list-like, dict | Values recognized as NA/NaN | None |
| encoding | str | File encoding, e.g., 'utf-8' | None |
Return Value
- Return Type:
pd.DataFrame - Returns a two-dimensional labeled data structure, i.e., a pandas DataFrame, on which various data analysis and processing operations can be performed.
Examples
Through the following examples, fully masterread_csv()the various usages.
Example 1: Reading a Local CSV File
First, create a simple CSV file, and then useread_csv()to read it.
Example
# Create a sample CSV file
# First write some test data
data = """name,age,city,salary
Tom,28,Beijing,8000
Jerry,35,Shanghai,12000
Mike,42,Guangzhou,15000
Lucy,26,Shenzhen,7000
"""
# Write the data to a file (in actual use, you can directly read an existing file)
with open('employees.csv', 'w', encoding='utf-8') as f:
f.write(data)
# Use read_csv to read the CSV file
# filepath_or_buffer: file path (required)
df = pd.read_csv('employees.csv')
# View the read result
print("The read DataFrame:")
print(df)
print("\nData types:")
print(df.dtypes)
print("\nColumn names:", df.columns.tolist())
Expected output:
读取的 DataFrame:
name age city salary
0 Tom 28 Beijing 8000
1 Jerry 35 Shanghai 12000
2 Mike 42 Guangzhou 15000
3 Lucy 26 Shenzhen 7000
数据类型:
name object
age int64
city object
salary int64
列名: ['name', 'age', 'city', 'salary']
Code explanation:
pd.read_csv('employees.csv')is the most basic usage; simply pass in the file path directly.- By default, the first row is automatically recognized as the column names (header='infer').
- pandas automatically infers the data type of each column: strings become object, and integers become int64.
- The returned DataFrame can be directly used with
print()to view, and subsequent data analysis can also be performed.
Example 2: Custom Column Names and Selecting Specific Columns
In real work, we may need to customize column names, or only read some columns to improve performance.
Example
# Create test data
data = """name,age,city,salary,department
Tom,28,Beijing,8000,IT
Jerry,35,Shanghai,12000,HR
Mike,42,Guangzhou,15000,Sales
Lucy,26,Shenzhen,7000,IT
"""
with open('employees2.csv', 'w', encoding='utf-8') as f:
f.write(data)
# Example 2a: Custom column names
# The names parameter is used to specify new column names, which will override the column names in the original file
df_custom = pd.read_csv('employees2.csv', names=['Name', 'Age', 'City', 'Salary', 'Department'], header=0)
print("DataFrame after custom column names:")
print(df_custom)
print()
# Example 2b: Only read the specified columns
# The usecols parameter can specify which columns to read, improving performance
df_partial = pd.read_csv('employees2.csv', usecols=['name', 'salary'])
print("Only reading some columns:")
print(df_partial)
print()
# Example 2c: Use index_col to specify the index column
df_indexed = pd.read_csv('employees2.csv', index_col='name')
print("Set name as the index:")
print(df_indexed)
Expected output:
自定义列名后的 DataFrame:
姓名 年龄 城市 薪资 部门
0 Tom 28 Beijing 8000 IT
1 Jerry 35 Shanghai 12000 HR
2 Mike 42 Guangzhou 15000 Sales
3 Lucy 26 Shenzhen 7000 IT
只读取部分列:
name salary
0 Tom 8000
1 Jerry 12000
2 Mike 15000
3 Lucy 7000
设置 name 为索引:
age city salary department
name
Tom 28 Beijing 8000 IT
Jerry 35 Shanghai 12000 HR
Mike 42 Guangzhou 15000 Sales
Lucy 26 Shenzhen 7000 IT
Code explanation:
namesThe parameter needs to match the number of columns in the data. If the file has a header row, you can setheader=0to use the original header.usecolsYou can pass a list of column names (recommended) or a list of column indices; the order of the returned columns matches the specified order.index_colSetting a column as the row index makes it easier to quickly query data by index later.
Example 3: Handling Special Data and Improving Missing Value Handling
CSV files may contain missing values, special delimiters, or situations where certain rows need to be skipped.
Example
# Create a CSV file containing special data
# Use semicolon as the delimiter, containing missing values and NA values
data = """name;age;city;salary
Tom;28;Beijing;8000
Jerry;;Shanghai;12000
Mike;42;Guangzhou;
Lucy;26;NA;7000
"""
with open('employees3.csv', 'w', encoding='utf-8') as f:
f.write(data)
# Example 3a: Read a semicolon-delimited file
df_semicolon = pd.read_csv('employees3.csv', sep=';')
print("Using semicolon delimiter:")
print(df_semicolon)
print("Missing value statistics:")
print(df_semicolon.isnull())
print()
# Example 3b: Specify which values are treated as missing values
df_na = pd.read_csv('employees3.csv', sep=';', na_values=['NA', 'missing'])
print("After custom NA values:")
print(df_na)
print()
# Example 3c: Skip rows and limit the number of rows read
# Assume the first few lines of the file are comments and can be skipped
data_with_comment = """# This is an employee data file
# Creation date: 2024-01-01
name;age;city;salary
Tom;28;Beijing;8000
Jerry;35;Shanghai;12000
Mike;42;Guangzhou;15000
Lucy;26;Shenzhen;7000
"""
with open('employees4.csv', 'w', encoding='utf-8') as f:
f.write(data_with_comment)
# skiprows skips the first two lines (comment lines)
df_skip = pd.read_csv('employees4.csv', sep=';', skiprows=2)
print("After skipping comment lines:")
print(df_skip)
print()
# nrows only reads the first 3 rows
df_nrows = pd.read_csv('employees4.csv', sep=';', skiprows=2, nrows=3)
print("Only reading the first 3 rows:")
print(df_nrows)
Expected output:
使用分号分隔符:
name age city salary
0 Tom 28.0 Beijing 8000.0
1 Jerry NaN Shanghai 12000.0
2 NaN 42.0 Guangzhou NaN
3 Lucy 26.0 NA 7000.0
缺失值统计:
name age city salary
0 False False False False
1 False True False False
2 True False False True
3 False False False False
自定义 NA 值后:
name age city salary
0 Tom 28.0 Beijing 8000.0
1 Jerry NaN Shanghai 12000.0
2 NaN 42.0 Guangzhou NaN
3 Lucy 26.0 NaN 7000.0
跳过注释行后:
name age city salary
0 Tom 28 Beijing 8000
1 Jerry 35 Shanghai 12000
42 Guangzhou 15000
2 Mike
3 Lucy 26 Shenzhen 7000
只读取前3行:
name age city salary
0 Tom 28 Beijing 800本地
1 Jerry 35 Shanghai 12000
2 Mike 42 Guangzhou 15000
Code explanation:
sepThe parameter can specify any delimiter, such as semicolon, tab, etc.- By default, empty strings, spaces, etc. are recognized as missing values. You can use the
na_valuesparameter to customize the values treated as NA. skiprowsYou can skip the specified number of rows at the beginning of the file, which is convenient for handling files with comments.nrowsLimits the number of rows read, suitable for chunked reading of large files.
Notes
- When reading large files, you can consider using the
chunksizeparameter for chunked reading to avoid insufficient memory. - When processing Chinese files, you need to correctly specify the
encodingparameter. Common encodings include 'utf-8', 'gbk', 'gb2312', etc. - If the CSV file has no header row, you need to set
header=None, and then use thenamesparameter to specify column names. - For non-standard CSV files, you may need to adjust
sep、quotecharand other parameters to parse them correctly.
Summary
read_csv()is the most basic and most important data reading function in pandas. It is powerful and supports advanced features such as multiple delimiters, custom column names, index settings, and missing value handling.
In practical data analysis work, masteringread_csv()the various parameter usages allows you to efficiently handle CSV files in various formats, laying a solid foundation for subsequent data cleaning and analysis. Readers are advised to practice more and become proficient in these common parameters.
Pandas Common Functions