Pandas pd.ead_csv() Function

Python math 模块Pandas Common Functions


read_csv()is the most commonly used data reading function in the pandas library, used to read data from CSV (Comma-Separated Values) files and create a DataFrame.

CSV files are a simple and widely used data exchange format that stores tabular data in plain text, where each row represents a record and fields are separated by commas.read_csv()It can intelligently parse CSV files, automatically recognize column names and data types, and handle various delimiter and encoding issues.


Basic Syntax and Parameters

Syntax Format

pandas.read_csv(filepath_or_buffer, sep=',', delimiter=None, header='infer',
                names=None, index_col=None, usecols=None, dtype=None,
                skiprows=None, nrows=None, na_values=None, ...)

Parameter Description

ParameterTypeDescriptionDefault Value
filepath_or_bufferstr, path object, or file-like objectThe path, URL, or file object of the CSV fileRequired
sepstrField delimiter, comma by default for CSV','
headerint, list of int, 'infer'Row number used as column names, 0 means the first row'infer'
nameslist-likeCustom list of column namesNone
index_colint, str, list of int, list of str, FalseColumn used as the row indexNone
usecolslist-like, callableOnly read the specified columnsNone
dtypedictData types for specified columns, e.g., {'a': np.float64}None
skiprowslist-like, intSkip the specified rowsNone
nrowsintOnly read the first n rowsNone
na_valuesscalar, str, list-like, dictValues recognized as NA/NaNNone
encodingstrFile encoding, e.g., 'utf-8'None

Return Value

  • Return Type:pd.DataFrame
  • Returns a two-dimensional labeled data structure, i.e., a pandas DataFrame, on which various data analysis and processing operations can be performed.

Examples

Through the following examples, fully masterread_csv()the various usages.

Example 1: Reading a Local CSV File

First, create a simple CSV file, and then useread_csv()to read it.

Example

import pandas as pd

# Create a sample CSV file
# First write some test data
data = """name,age,city,salary
Tom,28,Beijing,8000
Jerry,35,Shanghai,12000
Mike,42,Guangzhou,15000
Lucy,26,Shenzhen,7000
"""


# Write the data to a file (in actual use, you can directly read an existing file)
with open('employees.csv', 'w', encoding='utf-8') as f:
    f.write(data)

# Use read_csv to read the CSV file
# filepath_or_buffer: file path (required)
df = pd.read_csv('employees.csv')

# View the read result
print("The read DataFrame:")
print(df)
print("\nData types:")
print(df.dtypes)
print("\nColumn names:", df.columns.tolist())

Expected output:

读取的 DataFrame:
    name  age       city  salary
0    Tom   28    Beijing    8000
1  Jerry   35   Shanghai   12000
2   Mike   42  Guangzhou   15000
3   Lucy   26  Shenzhen    7000

数据类型:
name      object
age        int64
city      object
salary    int64

列名: ['name', 'age', 'city', 'salary']

Code explanation:

  • pd.read_csv('employees.csv')is the most basic usage; simply pass in the file path directly.
  • By default, the first row is automatically recognized as the column names (header='infer').
  • pandas automatically infers the data type of each column: strings become object, and integers become int64.
  • The returned DataFrame can be directly used withprint()to view, and subsequent data analysis can also be performed.

Example 2: Custom Column Names and Selecting Specific Columns

In real work, we may need to customize column names, or only read some columns to improve performance.

Example

import pandas as pd

# Create test data
data = """name,age,city,salary,department
Tom,28,Beijing,8000,IT
Jerry,35,Shanghai,12000,HR
Mike,42,Guangzhou,15000,Sales
Lucy,26,Shenzhen,7000,IT
"""


with open('employees2.csv', 'w', encoding='utf-8') as f:
    f.write(data)

# Example 2a: Custom column names
# The names parameter is used to specify new column names, which will override the column names in the original file
df_custom = pd.read_csv('employees2.csv', names=['Name', 'Age', 'City', 'Salary', 'Department'], header=0)
print("DataFrame after custom column names:")
print(df_custom)
print()

# Example 2b: Only read the specified columns
# The usecols parameter can specify which columns to read, improving performance
df_partial = pd.read_csv('employees2.csv', usecols=['name', 'salary'])
print("Only reading some columns:")
print(df_partial)
print()

# Example 2c: Use index_col to specify the index column
df_indexed = pd.read_csv('employees2.csv', index_col='name')
print("Set name as the index:")
print(df_indexed)

Expected output:

自定义列名后的 DataFrame:
    姓名  年龄       城市    薪资     部门
0   Tom   28    Beijing   8000      IT
1  Jerry   35   Shanghai  12000      HR
2   Mike   42  Guangzhou  15000   Sales
3   Lucy   26   Shenzhen   7000      IT

只读取部分列:
   name  salary
0    Tom    8000
1  Jerry   12000
2   Mike   15000
3    Lucy    7000

设置 name 为索引:
      age       city  salary department
name
Tom     28    Beijing    8000         IT
Jerry  35   Shanghai   12000         HR
Mike   42  Guangzhou   15000       Sales
Lucy   26   Shenzhen    7000         IT

Code explanation:

  • namesThe parameter needs to match the number of columns in the data. If the file has a header row, you can setheader=0to use the original header.
  • usecolsYou can pass a list of column names (recommended) or a list of column indices; the order of the returned columns matches the specified order.
  • index_colSetting a column as the row index makes it easier to quickly query data by index later.

Example 3: Handling Special Data and Improving Missing Value Handling

CSV files may contain missing values, special delimiters, or situations where certain rows need to be skipped.

Example

import pandas as pd

# Create a CSV file containing special data
# Use semicolon as the delimiter, containing missing values and NA values
data = """name;age;city;salary
Tom;28;Beijing;8000
Jerry;;Shanghai;12000
Mike;42;Guangzhou;
Lucy;26;NA;7000
"""


with open('employees3.csv', 'w', encoding='utf-8') as f:
    f.write(data)

# Example 3a: Read a semicolon-delimited file
df_semicolon = pd.read_csv('employees3.csv', sep=';')
print("Using semicolon delimiter:")
print(df_semicolon)
print("Missing value statistics:")
print(df_semicolon.isnull())
print()

# Example 3b: Specify which values are treated as missing values
df_na = pd.read_csv('employees3.csv', sep=';', na_values=['NA', 'missing'])
print("After custom NA values:")
print(df_na)
print()

# Example 3c: Skip rows and limit the number of rows read
# Assume the first few lines of the file are comments and can be skipped
data_with_comment = """# This is an employee data file
# Creation date: 2024-01-01
name;age;city;salary
Tom;28;Beijing;8000
Jerry;35;Shanghai;12000
Mike;42;Guangzhou;15000
Lucy;26;Shenzhen;7000
"""

with open('employees4.csv', 'w', encoding='utf-8') as f:
    f.write(data_with_comment)

# skiprows skips the first two lines (comment lines)
df_skip = pd.read_csv('employees4.csv', sep=';', skiprows=2)
print("After skipping comment lines:")
print(df_skip)
print()

# nrows only reads the first 3 rows
df_nrows = pd.read_csv('employees4.csv', sep=';', skiprows=2, nrows=3)
print("Only reading the first 3 rows:")
print(df_nrows)

Expected output:

使用分号分隔符:
    name  age       city   salary
0    Tom   28.0    Beijing   8000.0
1  Jerry   NaN   Shanghai  12000.0
2    NaN  42.0   Guangzhou      NaN
3  Lucy  26.0         NA    7000.0

缺失值统计:
    name    age   city  salary
0  False  False  False   False
1  False   True  False   False
2   True  False  False    True
3  False  False   False  False

自定义 NA 值后:
    name  age       city   salary
0    Tom  28.0    Beijing   8000.0
1  Jerry   NaN   Shanghai  12000.0
2   NaN  42.0   Guangzhou      NaN
3  Lucy  26.0         NaN   7000.0

跳过注释行后:
    name  age       city  salary
0    Tom   28   Beijing    8000
1  Jerry  35  Shanghai   12000
  42  Guangzhou   15000
2   Mike
3   Lucy   26   Shenzhen    7000

只读取前3行:
    name  age       city  salary
0    Tom   28   Beijing    800本地
1  Jerry  35   Shanghai   12000
2   Mike   42  Guangzhou   15000

Code explanation:

  • sepThe parameter can specify any delimiter, such as semicolon, tab, etc.
  • By default, empty strings, spaces, etc. are recognized as missing values. You can use thena_valuesparameter to customize the values treated as NA.
  • skiprowsYou can skip the specified number of rows at the beginning of the file, which is convenient for handling files with comments.
  • nrowsLimits the number of rows read, suitable for chunked reading of large files.

Notes

  • When reading large files, you can consider using thechunksizeparameter for chunked reading to avoid insufficient memory.
  • When processing Chinese files, you need to correctly specify theencodingparameter. Common encodings include 'utf-8', 'gbk', 'gb2312', etc.
  • If the CSV file has no header row, you need to setheader=None, and then use thenamesparameter to specify column names.
  • For non-standard CSV files, you may need to adjustsep、quotecharand other parameters to parse them correctly.

Summary

read_csv()is the most basic and most important data reading function in pandas. It is powerful and supports advanced features such as multiple delimiters, custom column names, index settings, and missing value handling.

In practical data analysis work, masteringread_csv()the various parameter usages allows you to efficiently handle CSV files in various formats, laying a solid foundation for subsequent data cleaning and analysis. Readers are advised to practice more and become proficient in these common parameters.

Python math 模块Pandas Common Functions

Other Extensions