Pandas df.astype() Function
df.astype()It is a function in Pandas for converting the data types of a DataFrame or Series.
During data processing, mismatched data types are a common problem. For example, numbers stored as strings, dates stored as text, etc.astype()It allows you to explicitly convert data to the required type, such as integer, float, string, date, etc., ensuring that the data can be correctly calculated and analyzed.
Basic Syntax and Parameters
astype()It is both a member function of DataFrame and a member function of Series, invoked via the dot operator.to call.
Syntax Format
DataFrame.astype(dtype, copy=True, errors='raise') Series.astype(dtype, copy=True, errors='raise')
Parameter Description
| Parameter | Type | Required? | Description | Default Value |
|---|---|---|---|---|
| dtype | Python dtype or numpy dtype | Required | Target data type. It can be a Python type (e.g.,int、str), a NumPy type (e.g.,np.int64、np.float32), or a Pandas type (e.g.,'int64'、'float64'、'category')。 |
None |
| copy | bool | Optional | If it isTrue, always return a new object; if it isFalse, modify the original object when possible. |
True |
| errors | str | Optional | Controls error handling.'raise'Indicates that an exception is raised when conversion fails;'ignore'Indicates that the original data is returned when conversion fails, without raising an exception. |
'raise' |
Return Value Description
- Returns a new DataFrame or Series in which the data types of all specified columns have been converted to the target type.
Examples
Let us thoroughly master through a series of examplesastype()the usage of.
Example 1: Convert the Data Type of a Single Column
Convert a column of a Series or DataFrame to the specified type.
Example
# Create a DataFrame where values are stored as strings
data = {
'Name': ['Zhang San', 'Li Si', 'Wang Wu', 'Zhao Liu'],
'age': ['25', '30', '35', '28'], # Store as string
'Salary': ['5000', '6000', '5500', '7000'] # Store as string
}
df = pd.DataFrame(data)
print("Original Data Type:")
print(df.dtypes)
print("=" * 50)
# Convert the "Age" column to integer type
df['age'] = df['age'].astype(int)
print(Data type of the age column after conversion:)
print(df.dtypes)
print("=" * 50)
# Convert the 'Salary' column to float type
df['Salary'] = df['Salary'].astype(float)
print(Data type of the salary column after conversion:)
print(df.dtypes)
print("=" * 50)
print("Converted Data:")
print(df)
Expected output:
原始数据类型:
姓名 object
年龄 object # 字符串类型
薪资 object # 字符串类型
==================================================
转换年龄列后的数据类型:
姓名 object
年龄 int64 # 整数类型
薪资 object
==================================================
转换薪资列后的数据类型:
姓名 object
年龄 int64
薪资 float64 # 浮点数类型
==================================================
转换后的数据:
姓名 年龄 薪资
0 张三 25 5000.0
1 李四 30 6000.0
2 王五 35 5500.0
3 赵六 28 7000.0
Code Explanation:
- In the original data, 'age' and 'salary' are both string types (
object)。 - Use
df['列名'].astype(int)to convert strings to integers. - After conversion, numerical calculations can be performed, such as sum, average, etc.
Example 2: Convert Data Types of an Entire DataFrame
You can specify data types for the entire DataFrame at once.
Example
# Create a DataFrame
data = {
'A': [1, 2, 3, 4],
'B': [1.5, 2.5, 3.5, 4.5],
'C': ['a', 'b', 'c', 'd']
}
df = pd.DataFrame(data)
print("Original Data Type:")
print(df.dtypes)
print("=" * 50)
# Convert column A to float64 and column B to int64
df_converted = df.astype({'A': 'float64', 'B': 'int64'})
print(Converted data types:)
print(df_converted.dtypes)
print("=" * 50)
print("Converted Data:")
print(df_converted)
Expected output:
原始数据类型:
A int64
B float64
C object
==================================================
转换后的数据类型:
A float64
B int64 # 小数部分被截断
C object
==================================================
转换后的数据:
A B C
0 1.0 2 a
1 2.0 2 b
2 3.0 3 c
3 4.0 4 d
Code Explanation:
- Use a dictionary
{'Column Name': 'Type'}You can specify the conversion types of multiple columns at once. - When converting the float values in column B to integers, the decimal part is truncated (2.5 becomes 2).
Example 3: Convert to Categorical Data Type
Categorical data type (category) can save memory, especially suitable for columns with limited values.
Example
# Create a DataFrame with many duplicate values
data = {
'city': ['Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen', 'Beijing', 'Shanghai', 'Guangzhou', 'Shenzhen'] * 1000,
Number: range(8000)
}
df = pd.DataFrame(data)
print("Original Data Type:")
print(df.dtypes)
print(fMemory usage: {df.memory_usage(deep=True).sum() / 1024:.2f} KB)
print("=" * 50)
# Convert the "City" column to categorical type
df['city'] = df['city'].astype('category')
print(Converted data types:)
print(df.dtypes)
print(fMemory usage: {df.memory_usage(deep=True).sum() / 1024:.2f} KB)
Expected output:
原始数据类型: 城市 object 编号 int64 内存使用: ~456 KB ================================================== 转换后的数据类型: 城市 category 编号 int64 内存使用: ~120 KB # 内存使用大幅减少
Code Explanation:
- The categorical data type assigns an integer identifier to each unique value, saving a lot of memory compared to storing full strings.
- The original data has 8000 rows and uses about 456 KB of memory; after conversion it is about 120 KB, saving about 75% of memory.
- This method is especially suitable for categorical variables with few values.
Example 4: Convert to Datetime Type
Convert strings to datetime type so that date-related operations can be performed.
Example
# Create a DataFrame with date strings
data = {
Date: ['2024-01-01', '2024-01-02', '2024-01-03', '2024-01-04'],
'Sales': [100, 150, 120, 180]
}
df = pd.DataFrame(data)
print("Original Data Type:")
print(df.dtypes)
print("=" * 50)
# Convert the "date" column to datetime type
df[Date] = pd.to_datetime(df[Date])
print(Converted data types:)
print(df.dtypes)
print("=" * 50)
print("Converted Data:")
print(df)
# Now you can perform date-related operations
print(nExtract month:)
print(df[Date].dt.month)
Expected output:
原始数据类型:
日期 object # 字符串类型
销量 int64
==================================================
转换后的数据类型:
日期 datetime64[ns] # 日期时间类型
销量 int64
==================================================
转换后的数据:
日期 销量
0 2024-01-01 100
1 2024-01-02 150
2 2024-01-03 120
3 2024-01-04 180
提取月份:
0 1
1 1
2 1
3 1
Name: 日期, dtype: int64
Code Explanation:
- Use
pd.to_datetime()to convert strings to datetime type. - After conversion, you can use
.dtaccessor to perform date operations, such as extracting month, day of the week, etc. - Note: Although you can also use
astype('datetime64[ns]'), butpd.to_datetime()is more powerful and flexible.
Example 5: Handling Conversion Errors
Useerrors='ignore'parameter to handle conversion failures.
Example
# Create a Series containing values that cannot be converted to integers
s = pd.Series(['1', '2', 'hello', '4', 'world'])
print(“Raw data:”)
print(s)
print("=" * 50)
# Try conversion. Using errors='raise' (default) will raise an exception.
try:
s_converted = s.astype(int)
except ValueError as e:
print(fConversion failed: {e})
print("=" * 50)
# Use errors='ignore', return the original data when conversion fails
s_converted = s.astype(int, errors='ignore')
print("Data converted using errors='ignore':")
print(s_converted)
print("=" * 50)
# Use errors='coerce' to fill with NaN when conversion fails
s_converted2 = s.astype(float, errors='coerce')
print("Data converted using errors='coerce':")
print(s_converted2)
Expected output:
原始数据: 0 1 1 2 2 hello # 无法转换为整数 3 4 4 world # 无法转换为整数 ================================================== 转换失败: invalid literal for int() with base 10: 'hello' ================================================== 使用 errors='ignore' 转换后的数据: 0 1 1 2 2 hello # 保持原值 3 4 4 world # 保持原值 ================================================== 使用 errors='coerce' 转换后的数据: 0 1.0 1 2.0 2 NaN # 无法转换,设为 NaN 3 4.0 4 NaN # 无法转换,设为 NaN
Code Explanation:
- When a string contains a value that cannot be converted to a number, by default it raises
ValueErroran exception. errors='ignore'will keep the original data unchanged.errors='coerce'will set the non-convertible value toNaN, which is a common method for handling dirty data.
Example 6: Convert to Boolean Type
Convert data to boolean type, commonly used in conditional filtering and logical operations.
Example
# Create a DataFrame containing 0 and 1
data = {
'Is Member': [0, 1, 1, 0, 1],
'Is Active': ['Yes', 'No', 'Yes', 'Yes', 'No']
}
df = pd.DataFrame(data)
print("Raw data:")
print(df)
print("Raw data types:")
print(df.dtypes)
print("=" * 50)
# Convert 0/1 to boolean type
df['Is Member'] = df['Is Member'].astype(bool)
print("Data types after converting the Is Member column:")
print(df.dtypes)
print(df)
print("=" * 50)
# Convert "Yes"/"No" to boolean type
df['Is Active'] = df['Is Active'].map({'Yes': True, 'No': False})
print("Data after converting the Is Active column:")
print(df)
print(df.dtypes)
Expected running results:
原始数据: 是否会员 是否激活 0 0 是 1 1 否 2 1 是 3 0 是 4 1 否 ================================================== 转换是否会员列后的数据类型: 是否会员 bool 是否激活 object ================================================== 转换是否激活列后的数据: 是否会员 是否激活 0 True True 1 True False 2 True True 3 False True 4 True False
Code analysis:
- 0 can be converted to
False, 1 converted toTrue。 - For strings like "Yes"/"No", you need to use
map()to convert. - After conversion, boolean indexing can be used for fast filtering.
Notes
astype()By default, a new object is returned, and the original data is not modified. If you want to modify in place, you can usecopy=Falseparameter, but this may affect the original data.- When converting floating-point numbers to integers, the decimal part is truncated, not rounded.
- When converting strings to numbers, make sure the string is in a valid numeric format, otherwise an exception will be thrown.
- Using
errors='coerce'is the recommended method for handling dirty data, and can convert invalid values toNaN, then usefillna()to fill. - For columns with many duplicate values, converting to a categorical type can significantly reduce memory usage.
Other Extensions