Pandas pd.notna() Function
pd.notna()is a function in the Pandas library used fordetecting missing values. It checks each element in the input object to determine whether it is a non-missing value, and returns a result consisting of boolean values (True/False).
It is one of the most commonly used functions in data cleaning, helping us quickly identify valid values in the data and prepare for subsequent data processing and analysis.
Term Explanation: not"not" is a negation,na"na" is an abbreviation for "not available", so together it means "not a missing value".
Basic Syntax and Parameters
pd.notna()is a top-level function of the Pandas library and can be called directly viapd.notna(), or through the Series or DataFrame object's.notna()method.
Syntax Format
pd.notna(obj)
Parameter Description
- Parameter:
obj- Type: Any Python object, such as Series, DataFrame, list, array, scalar value, etc.
- Description: The object to check for non-missing values. Missing values in Pandas include
None、NaN、NaT(missing time type) andpandas.NA。
Function Description
- Return Value: Returns a boolean object with the same shape as the input object. If it is a Series or DataFrame, returns an object of the same type; if it is a scalar, returns a single boolean value.
- Effect: Returns True for non-missing values
True, and returns False for missing valuesFalse。
Examples
Let us thoroughly master the usage of pd.notna()pd.notna()through a series of examples from simple to complex.
Example 1: Basic Usage - Detecting Missing Values in Scalars and Lists
Example
import numpy as np
# 1. Check whether scalar values are non-missing
print("=== Scalar Detection ===")
print(f"pd.notna(10): {pd.notna(10)}") # Normal numeric value -> True
print(f"pd.notna('example'): {pd.notna('example')}") # Normal string -> True
print(f"pd.notna(None): {pd.notna(None)}") # None -> False
print(f"pd.notna(np.nan): {pd.notna(np.nan)}") # np.nan -> False
# 2. Detect missing values in a list
print("\n=== List Detection ===")
data_list = [1, 2, np.nan, 'example', None, 5]
result = pd.notna(data_list)
print(f"Original list: {data_list}")
print(f"Detection result: {result.tolist()}") # Convert to list for easier viewing
Expected output:
=== 标量检测 ===
pd.notna(10): True
pd.notna('example'): True
pd.notna(None): False
pd.notna(np.nan): False
=== 列表检测 ===
原始列表: [1, 2, nan, 'example', None, 5]
检测结果: [True, False, True, False, True]
Code explanation:
pd.notna(10)andpd.notna('example')Detects normal numeric values and strings, returns TrueTrue。pd.notna(None)andpd.notna(np.nan)Detects missing value markers of Python and NumPy, returns FalseFalse。- Using pd.notna() on a list
pd.notna()returns a boolean list, where missing value positions are FalseFalse。
Example 2: Detecting Missing Values in a Series
When processing Series data,pd.notna()pd.notna() can quickly locate valid data.
Example
import numpy as np
# Create a Series containing missing values
s = pd.Series([1, 2, np.nan, 4, None, 'example', np.nan])
print("=== Original Series ===")
print(s)
print(f"\nType: {type(s)}")
# Use pd.notna() to detect
print("\n=== pd.notna() detection result ===)
result = pd.notna(s)
print(result)
# Use the Series' notna() method (equivalent effect)
print("\n=== s.notna() method detection ===)
print(s.notna())
# Filter out non-missing values
print("\n=== Filter non-missing values ===)
print(s[s.notna()])
Expected output:
=== 原始 Series === 0 1 1 2 2 NaN 3 4 4 None 5 example 6 NaN dtype: object === pd.notna() 检测结果 === 0 True 1 True 2 False 3 True 4 False 5 True 6 False dtype: bool === s.notna() 方法检测 === 0 True 1 True 2 False 3 True 4 False 5 True 6 False dtype: bool === 筛选非缺失值 === 0 1 1 2 3 4 5 example dtype: object
Code explanation:
pd.notna(s)ands.notna()Both can detect missing values in a Series and return a boolean Series.- Using boolean indexing
s[s.notna()]can quickly filter out all non-missing elements, which is very useful in data cleaning.
Example 3: Detecting Missing Values in a DataFrame
When processing tabular data,pd.notna()pd.notna() can quickly reveal the completeness of the data.
Example
import numpy as np
# Create a DataFrame containing missing values
df = pd.DataFrame({
'name': ['Alice', 'Bob', None, 'Diana'],
'age': [25, np.nan, 30, 28],
'score': [85, 90, np.nan, 95]
})
print("=== Original DataFrame ===")
print(df)
# Detect the entire DataFrame
print("\n=== pd.notna() detection result ===)
print(pd.notna(df))
# Count missing values by column
print("\n=== Number of non-missing values per column ===)
print(df.notna().sum())
# Count missing values by row
print("\n=== Number of non-missing values per row ===)
print(df.notna().sum(axis=1))
# Calculate missing value ratio
print("\n=== Missing value ratio ===)
missing_ratio = df.isna().mean()
print(missing_ratio)
Expected output:
=== 原始 DataFrame ===
name age score
0 Alice 25 85.0
1 Bob NaN 90.0
2 None 30 NaN
3 Diana 28 95.0
=== pd.notna() 检测结果 ===
name age score
0 True True True
1 True False True
2 False True False
3 True True True
=== 每列非缺失值数量 ===
name 3
age 3
score 3
=== 每行非缺失值数量 ===
0 3
1 2
2 1
3 3
=== 缺失值比例 ===
name 0.25
age 0.25
score 0.25
Code explanation:
pd.notna(df)Returns a boolean DataFrame with the same shape as the original DataFrame, where each position indicates whether it is a non-missing value.df.notna().sum()Counts the number of non-missing values in each column.df.isna().mean()Calculates the proportion of missing values in each column to help assess data quality.
Other ExtensionsTip:
pd.notna()andpd.isna()are opposite to each other.pd.notna()Where pd.notna() returns True,pd.isna()pd.isna() returns False, and vice versa.
Pandas Common Functions