Pandas pd.to-numeric() Function
pd.to_numeric()is used in the Pandas library toconvert data to numeric typesfunction. It can convert data in various formats (such as strings, mixed-type columns) to numeric types such as integers and floats.
This is a commonly used function in data cleaning, especially when processing data imported from files or databases, where it is often necessary to convert columns that look numeric but are actually strings.
Word Definition: to_numericIt means "convert to numeric", that is, converting data from other types to numeric types.
Basic Syntax and Parameters
pd.to_numeric()is a top-level function of the Pandas library, used to convert input to numeric types.
Syntax Format
pd.to_numeric(arg, errors='raise', downcast=None)
Parameter Description
- Parameter:
arg- Type: Series, list, array, or dictionary-like object.
- Description: The data to be converted to numeric types. Usually a Series.
- Parameter:
errors- Type: String ('raise', 'coerce', 'ignore').
- Description: Error handling mode.
'raise'(Default) Throws an exception when encountering a value that cannot be converted;'coerce'Sets unconvertible values to NaN;'ignore'Ignores errors and returns the original data.
- Parameter:
downcast- Type: String ('integer', 'signed', 'unsigned', 'float') or None.
Function Description
- Return value: Returns a numeric Series.
- Effect: Converts the type of the input data to numeric (int64 or float64).
Examples
Let's thoroughly master through a series of examples from simple to complex,pd.to_numeric()the usage of.
Example 1: Basic Usage - Converting Strings to Numeric Values
Example
# 1. Create a Series containing numeric strings
s = pd.Series(['10', '20', '30', '40', '50'])
print("=== Original Series (type:", s.dtype, ")===")
print(s)
# 2. Use pd.to_numeric() to convert to numeric type
result = pd.to_numeric(s)
print("\n=== After pd.to_numeric() conversion (type:", result.dtype, ")===")
print(result)
Expected output:
=== 原始 Series(类型: object )=== 0 10 1 20 2 30 3 40 5 50 dtype: object === pd.to_numeric() 转换后(类型: int64 )=== 0 10 1 20 2 30 3 40 4 50 dtype: int64
Code analysis:
- The original data is of string type (object), so numeric operations cannot be performed.
pd.to_numeric()After converting it to integer type (int64), mathematical operations can be performed.
Example 2: Handling Strings Containing Non-numeric Values
Usingerrorsthe parameter can flexibly handle values that cannot be converted.
Example
import numpy as np
# 1. Create a Series containing non-numeric strings
s = pd.Series(['10', '20', 'abc', '40', 'example'])
print("=== Series containing non-numeric values ===")
print(s)
# 2. errors='raise' (default) - throws an exception when encountering an unconvertible value
print("\n=== errors='raise' (default) ===)
try:
result = pd.to_numeric(s, errors='raise')
except Exception as e:
print(f"Exception: {e}")
# 3. errors='coerce' - set unconvertible values to NaN
print("\n=== errors='coerce' ===")
result_coerce = pd.to_numeric(s, errors='coerce')
print(result_coerce)
# 4. errors='ignore' - ignore errors and return the original data
print("\n=== errors='ignore' ===")
result_ignore = pd.to_numeric(s, errors='ignore')
print(result_ignore)
print(f"Type: {result_ignore.dtype}")
Expected output:
=== 包含非数值的 Series === 0 10 1 20 2 abc 3 40 4 example 包含了 'abc', 'example' 等无法转换的字符串 === errors='raise'(默认)=== 异常: Unable to convert string to float explicitly === errors='coerce' === 0 10.0 1 20.0 2 NaN 3 40.0 4 NaN dtype: float64 === errors='ignore' === 0 10 1 20 2 abc 3 40 4 example Type: object
Code analysis:
errors='coerce'Very practical: it replaces unconvertible values with NaN (missing values) while retaining convertible values.errors='ignore'Keeps the original data unchanged, suitable when you only want to try the conversion without changing the data.
Example 3: Handling Mixed Numeric Values and Missing Values
In real data, it is often necessary to handle columns that contain a mix of missing values and numeric values.
Example
import numpy as np
# 1. Create a Series containing missing values and numeric strings
s = pd.Series(['100', '200', None, 'N/A', '400', '', '500'])
print("=== Series containing missing values and non-numeric values ===")
print(s)
# 2. Convert using errors='coerce'
print("\n=== Converted using errors='coerce' ===)
result = pd.to_numeric(s, errors='coerce')
print(result)
# 3. Check which values were converted to NaN
print("\n=== Identifying NaN values ===)
print(f"NaN positions: {result.isna().tolist()}")
# 4. Fill NaN with 0 or delete
print("\n=== Fill NaN with 0 ===)
print(result.fillna(0))
Expected output:
=== 包含缺失值和数值字符串的 Series === 0 100 1 200 2 None 3 N/A 4 400 5 (空字符串) 6 字符串 '500' dtype: object === 使用 errors='coerce' 转换 === 0 100.0 1 200.0 даль 2 NaN 3. NaN 4 400.0 5 NaN 6 NaN dtype: float64
Code analysis:
errors='coerce'It can handle None, empty strings, N/A, and other forms representing missing values, converting them all uniformly to NaN.- The converted type becomes float64 (because NaN needs to be represented).
Example 4: Using the downcast Parameter to Optimize Memory
For large datasets, you can usedowncastthe parameter to reduce memory usage.
Example
import numpy as np
# 1. Create a Series of large integers
s = pd.Series([1, 2, 3, 4, 5] * 100000)
print("=== Original type:", s.dtype)
print("=== Original memory:", s.memory_usage(deep=True), "bytes")
# 2. Convert to numeric without specifying downcast
result_default = pd.to_numeric(s)
print("\n=== Type after default conversion:", result_default.dtype)
print("=== Default memory:", result_default.memory_usage(deep=True), "bytes")
# 3. Downcast to a smaller integer type
result_signed = pd.to_numeric(s, downcast='signed')
print("\n=== downcast='signed' type:", result_signed.dtype)
print("=== downcast='signed' memory:", result_signed.memory_usage(deep=True), "bytes")
# 4. Downcast to floating-point numbers
float_data = pd.Series([1.5, 2.5, 3.5] * 100000)
result_float = pd.to_numeric(float_data, downcast='float')
print("\n=== downcast='float' type:", result_float.dtype)
print("=== downcast='float' memory:", result_float.memory_usage(deep=True), "bytes")
Expected output:
=== 原始类型: int64 === 原始内存: 2800000 bytes === 默认转换后类型: int64 内存节省:2.8 MB -> 1.4 MB === downcast='signed' 类型: int8/int16/int32 内存节省约 50% === downcast='float' 类型: float32 内存节省约 50%
Code analysis:
- For large datasets,
downcastit can significantly reduce memory usage. downcast='signed'Attempts to convert integers to the smallest signed integer type.downcast='float'Converts floating-point numbers from float64 to float32.
Other ExtensionsTip:When processing large-scale data,
pd.to_numeric()combiningerrors='coerce'anddowncastthe parameter can efficiently convert mixed data to numeric types and optimize memory.
Pandas Common Functions