Pandas Data Types (dtype and Type Conversion)
Pandas provides a rich data type system. Correctly understanding and using data types is the foundation of efficient data analysis. This section details Pandas' data type system, type inference, and type conversion methods.
Pandas Data Type Overview
| dtype | Description | Python Type | Example |
|---|---|---|---|
int64 |
64-bit Integer | int | 1, 2, 100 |
float64 |
64-bit Float | float | 1.5, 3.14 |
object |
String or Mixed Type | str | "hello" |
bool |
Boolean | bool | True, False |
datetime64[ns] |
Datetime | datetime | 2024-01-01 |
timedelta64[ns] |
Timedelta | timedelta | 1 days |
category |
Categorical Type | - | Finite Set |
Example
import numpy as np
# Create a DataFrame with various data types
df = pd.DataFrame({
"Integer": [1, 2, 3],
"Float": [1.5, 2.5, 3.5],
"String": ["a", "b", "c"],
"Boolean": [True, False, True],
"Date": pd.date_range("2024-01-01", periods=3)
})
print("Column data types:")
print(df.dtypes)
Type Inference and Specification
Automatic Type Inference
Example
# Type inference when reading CSV
# pandas attempts to infer the most suitable type for each column
df = pd.read_csv("data.csv")
# Or use the dtype parameter to explicitly specify types
df = pd.read_csv("data.csv", dtype={
"Age": "int32", # Specify as 32-bit integer to save memory
"Salary": "float32", # Specify as 32-bit float
"Name": "string" # Use PyArrow string
})
Specifying Types at Creation
Example
import numpy as np
# Specify data type to create Series
s = pd.Series([1, 2, 3], dtype="int8") # Use a smaller integer type
print(f"int8 type: {s.dtype}")
s = pd.Series([1.5, 2.5, 3.5], dtype="float32") # Use float32
print(f"float32 type: {s.dtype}")
# Use numpy type
s = pd.Series([1, 2, 3], dtype=np.int8)
print(f"np.int8 type: {s.dtype}")
Type Conversion
Converting with astype
Example
import numpy as np
# Create sample data
df = pd.DataFrame({
"Integer": [1, 2, 3],
"Float": [1.5, 2.5, 3.5],
"String": ["1", "2", "3"],
"Boolean": [1, 0, 1]
})
print("Original types:")
print(df.dtypes)
print()
# Convert to string
df["Integer_str"] = df["Integer"].astype(str)
print("After converting to string:")
print(df.dtypes)
# Convert string to numeric
df["String_int"] = df["String"].astype(int)
print("\n"String to integer:")
print(df.dtypes)
# Convert numeric to boolean (non-zero is True)
df["Boolean_int"] = df["Boolean"].astype(bool)
print("\n"Integer to boolean:")
print(df.dtypes)
# Convert float to integer (truncation)
df["Float_int"] = df["Float"].astype(int)
print("\n"Float to integer:")
print(df)
Converting with pd.to_numeric
Example
import numpy as np
# Handle numeric strings with special characters
s = pd.Series(["$1,000", "$2,500", "$3,200"])
# Clean and convert
s_cleaned = s.str.replace("$", "", regex=False).str.replace(",", "", regex=False)
s_numeric = pd.to_numeric(s_cleaned)
print("String to numeric:")
print(s_numeric)
print(f"Type: {s_numeric.dtype}")
print()
# Handle missing values
s_with_na = pd.Series(["1", "2", "NA", "4"])
s_numeric = pd.to_numeric(s_with_na, errors="coerce") # Convert invalid values to NaN
print("Handling missing values:")
print(s_numeric)
Converting Dates with pd.to_datetime
Example
# Convert various date formats
dates = ["2024-01-01", "2024/01/02", "01/03/2024", "20240104"]
# Convert dates
dt = pd.to_datetime(dates, errors="coerce")
print("Date conversion result:")
print(dt)
print(f"Type: {dt.dtype}")
print()
# Specify date format
dates2 = ["20240101", "20240102", "20240103"]
dt2 = pd.to_datetime(dates2, format="%Y%m%d")
print("Conversion with specified format:")
print(dt2)
Memory Optimization
Choosing appropriate data types can significantly reduce memory usage.
Choosing Integer Types
Example
import numpy as np
# Create a Series with a large amount of data
s = pd.Series(np.random.randint(0, 100, 1000000))
# Default int64
print(f"int64 memory: {s.dtype} -> {s.memory_usage(deep=True) / 1024 / 1024:.2f} MB")
# Convert to a smaller type
s_int8 = s.astype("int8")
print(f"int8 memory: {s_int8.dtype} -> {s_int8.memory_usage(deep=True) / 1024 / 1024:.2f} MB")
# Choose an appropriate type based on the data range
# np.iinfo to view integer range
print(f"\n"int8 range: {np.iinfo('int8').min} to {np.iinfo('int8').max}")
print(f"int16 range: {np.iinfo('int16').min} to {np.iinfo('int16').max}")
print(f"int32 range: {np.iinfo('int32').min} to {np.iinfo('int32').max}")
Using the category Type
Example
import numpy as np
# Create a Series with many duplicate values
s = pd.Series(np.random.choice(["Beijing", "Shanghai", "Guangzhou", "Shenzhen"], 1000000))
# String type
print(f"object type memory: {s.dtype} -> {s.memory_usage(deep=True) / 1024 / 1024:.2f} MB")
# Convert to category type
s_cat = s.astype("category")
print(f"category type memory: {s_cat.dtype} -> {s_cat.memory_usage(deep=True) / 1024 / 1024:.2f} MB")
# Disadvantage: some operations may become slower
print("\n"category info:")
print(s_cat.cat.categories)
Type Checking and Validation
Example
import numpy as np
df = pd.DataFrame({
"Integer": [1, 2, 3],
"Float": [1.5, 2.5, 3.5],
"String": ["a", "b", "c"]
})
# Check data types
print("Check if it is a numeric type:")
print(df.dtypes)
# Use is_integer_dtype
print(f"\n"Integer column: {pd.api.types.is_integer_dtype(df['Integer'])}")
print(f"Float column: {pd.api.types.is_float_dtype(df['Float'])}")
print(f"String column: {pd.api.types.is_object_dtype(df['String'])}")
# Check if it can be converted to numeric
print(f"\n"Can be converted to numeric: {pd.api.types.is_numeric_dtype(df['Integer'])}")
print(f"Is datetime: {pd.api.types.is_datetime64_any_dtype(pd.Series(pd.date_range('2024', periods=3)))}")
Common Issues
1. Float values appear in integer columns
When mixed into a DataFrame, it automatically converts to float64. UseNullable IntegerThe type can preserve integer null values.
2. Inconsistent date formats
Usepd.to_datetimeand specifyformatparameter to handle non-standard formats.
3. Type error when reading CSV
Usedtypethe parameter to explicitly specify the type, or useconvertersto convert.
Other extensionsFor large datasets, prioritize using smaller data types (such as int8, int16, float32) and category types to reduce memory usage.