Pandas Series.str.contains() Function

Pandas 常用函数Commonly Used Pandas Functions


Series.str.contains()is a function in Pandas used to check whether a string contains a specified substring.

In data processing, we often need to filter data based on text content, such as finding records containing specific keywords or filtering out text that matches a certain pattern.contains()The function can check whether each string element contains a specified substring or regular expression pattern.

Word Meaning:containsmeans 'to contain', indicating whether the string contains specified content.


Basic Syntax and Parameters

str.contains()is the string accessor method of Series, so you need to first have a Series containing strings, and then use the.straccessor to call it.

Syntax Format

Series.str.contains(pat, case=True, regex=True, na=None)

Parameter Description

Parameter Type Required Description Default Value
pat str Required The pattern to search for, which can be an ordinary string or a regular expression. -
case bool Optional Whether to distinguish case. Defaults to True (case-sensitive). True
regex bool Optional Whether to treat the pat parameter as a regular expression. Defaults to True. True
na object Optional The value to return when the element is NaN. Defaults to None (returns NaN). None

Function Description

  • Return Value: returns a boolean Series indicating whether each element contains the specified pattern.
  • Effect: checks each string element in the Series and returns True or False.
  • Note: uses regular expression matching by default, which can match more complex patterns.

Examples

Let's thoroughly master, through a series of examples from simple to complex,str.contains()the usage of it.

Example 1: Basic Usage - Check if a Substring Is Contained

Example

import pandas as pd

# Create a Series containing text
s = pd.Series(['apple', 'banana', 'grape', 'pineapple', 'orange'])

# Check whether it contains 'ap'
result = s.str.contains('ap')

print("Original Series:")
print(s)
print("nContains 'ap':")
print(result)

Output result:

原始 Series:
0       apple
1      banana
2       grape
3    pineapple
4      orange
dtype: object

是否包含 'ap':
0       True
1      False
2       True
3       True
4      False

Code explanation:

  1. s.str.contains('ap')Check whether each string contains the substring 'ap'.
  2. 'apple' contains 'ap', returns True.
  3. 'banana' does not contain 'ap', returns False.
  4. 'grape' contains 'ap', returns True.
  5. 'pineapple' contains 'ap', returns True.

Example 2: Using Regular Expression Matching

contains()Regular expressions are used by default, which can match more complex patterns.

Example

import pandas as pd

# Create a Series containing text
s = pd.Series(['hello123', 'world456', 'example789', 'python', 'test123'])

# Use a regular expression to match strings that start with a letter and are followed by digits
result = s.str.contains(r'^[a-zA-Z]+d+$')

print("Original Series:")
print(s)
print("nMatch the pattern starting with a letter and ending with a digit:")
print(result)

Output result:

原始 Series:
0      hello123
1      world456
2     example789
3       python
4      test123
dtype: object

匹配字母开头+数字结尾的模式:
0       True
1       True
2       True
3      False
4       True

Code explanation:

  • r'^[a-zA-Z]+d+$'It is a regular expression that matches strings starting with a letter and ending with a digit.
  • ^indicates the beginning of the string,$indicates the end of the string.
  • 'python' does not contain digits, so it does not match.

Example 3: Case-Insensitive Matching

By settingcase=Falsecase-insensitive matching can be achieved.

Example

import pandas as pd

# Create a Series containing text in different cases
s = pd.Series(['Apple', 'APPLE', 'apple', 'Banana', 'APPLE'])

# Case-sensitive matching
result_case = s.str.contains('APPLE')
# Case-insensitive matching
result_nocase = s.str.contains('APPLE', case=False)

print("Original Series:")
print(s)
print("nCase-sensitive match 'APPLE':")
print(result_case)
print("nCase-insensitive match 'APPLE':")
print(result_nocase)

Output result:

原始 Series:
0      Apple
1     APPLE
2     apple
3     Banana
4     APPLE
dtype: object

区分大小写匹配 'APPLE':
0      False
1       True
2      False
3      False
4       True

不区分大小写匹配 'APPLE':
0       True
1       True
2       True
3      False
4       True

Code explanation:

  • case=True(default) Case-sensitive, only matches exactly 'APPLE'.
  • case=FalseCase-insensitive, 'Apple', 'APPLE', and 'apple' all match.

Example 4: Handling Missing Values

Throughnathe parameter, you can specify how to handle NaN values.

Example

import pandas as pd
import numpy as np

# Create a Series containing NaN values
s = pd.Series(['apple', 'banana', np.nan, 'grape', None])

# Default handling of NaN (returns NaN)
result_default = s.str.contains('ap')
# Specify that NaN returns False
result_na_false = s.str.contains('ap', na=False)

print("Original Series:")
print(s)
print("nDefault handling (returns NaN):")
print(result_default)
print("nTreat NaN as False:")
print(result_na_false)

Output result:

原始 Series:
0       apple
1      banana
2        NaN
3       grape
4       None
dtype: object

默认处理(返回 NaN):
0       True
1      False
2        NaN
3       True
4        NaN

将 NaN 视为 False:
0       True
1      False
2      False
3       True
4      False

Code explanation:

  • By default, NaN and None return NaN.
  • Settingna=Falsecan treat NaN as not matching.
  • In actual data processing, this is useful and can avoid filtering issues caused by NaN.

Example 5: Filtering Data

contains()It is often combined with boolean indexing to filter data.

Example

import pandas as pd

# Create a simulated product data Series
products = pd.Series([
    'iPhone 14 Pro',
    'Samsung Galaxy S23',
    'iPhone 13',
    'Google Pixel 7',
    'iPad Pro',
    'MacBook Air',
    'Dell XPS 15'
])

# Filter products containing 'iPhone'
iphone_products = products[products.str.contains('iPhone')]

# Filter products that do not contain 'i' (case-sensitive)
no_i_products = products[~products.str.contains('i')]

print("All products:")
print(products)
print("nProducts containing 'iPhone':")
print(iphone_products)
print("nProducts not containing uppercase 'I':")
print(no_i_products)

Output result:

所有产品:
0       iPhone 14 Pro
1    Samsung Galaxy S23
2          iPhone 13
3    Google Pixel 7
4          iPad Pro
5      MacBook Air
6        Dell XPS 15
dtype: object

包含 'iPhone' 的产品:
0    iPhone 14 Pro
2          iPhone 13
dtype: object

不包含大写 'I' 的产品:
6    Dell XPS 15

Code explanation:

  • products[products.str.contains('iPhone')]Filter out products containing 'iPhone'.
  • ~products.str.contains('i')Use~negation to filter out products that do not contain 'i'.
  • This is a common operation in data filtering.

Notes

  • str.contains()Regular expressions are used by default (regex=True)。
  • If you need to match ordinary strings (not as regular expressions), you can setregex=False。
  • By default, it is case-sensitive; usecase=Falseto make it case-insensitive.
  • By default, NaN values return NaN; usenaparameter to specify the return value.
  • This function returns a boolean Series, which can be directly used for boolean indexing to filter data.

Pandas 常用函数Commonly Used Pandas Functions

Other Extensions