Pandas Series.str.replace() Function

Pandas 常用函数Pandas Common Functions


Series.str.replace()It is a function in Pandas used to replace specified content in strings.

In data processing, we often need to perform text replacement operations, such as correcting erroneous text, standardizing data formats, removing unwanted characters, etc.replace()The function can flexibly replace specified content in strings and supports regular expression matching.

Word Meaning:replaceIt means "replace," indicating replacing old content with new content.


Basic Syntax and Parameters

str.replace()It is the string accessor method of Series, so you first need a Series containing strings, and then use the.straccessor to call it.

Syntax Format

Series.str.replace(pat, repl, regex=False)

Parameter Description

Parameter Type Required Description Default Value
pat str or regex Required The pattern to replace, can be a plain string or a regular expression (when regex=True). -
repl str or callable Required The replacement content, can be a string or a callable function (for advanced replacement). -
regex bool Optional Whether to treat the pat parameter as a regular expression. Defaults to False. False

Function Description

  • Return Value: Returns a new Series after replacement.
  • Effect: Replaces the matched content in each string element of the Series with new content.
  • Note: By default, plain string matching is used. If regular expression matching is needed, you need to setregex=True。

Examples

Let's go through a series of examples from simple to complex to thoroughly masterstr.replace()its usage.

Example 1: Basic Usage - Simple String Replacement

Example

import pandas as pd

# Create a Series containing text
s = pd.Series(['apple', 'banana', 'apricot', 'pineapple'])

# Replace 'ap' with 'XX'
result = s.str.replace('ap', 'XX')

print("Original Series:")
print(s)
print("nResult after replacement:")
print(result)

Output result:

原始 Series:
0        apple
1       banana
2       apricot
3    pineapple
dtype: object

替换后的结果:
0       XXle
1       banana
2       XXricot
3    XXneXXple

Code explanation:

  1. s.str.replace('ap', 'XX')Replaces all occurrences of 'ap' with 'XX'.
  2. By default, only the first match is replaced (if it appears multiple times in the string, only the first one is replaced).
  3. You can replace all matches by adding a regular expression flag.

Example 2: Using Regular Expressions to Replace All Matches

When you need to replace all matches, you can use regular expressions.

Example

import pandas as pd

# Create a Series containing text
s = pd.Series(['apple', 'banana', 'apricot', 'pineapple'])

# Use a regular expression to replace all 'ap' with 'XX'
result = s.str.replace('ap', 'XX', regex=True)

print("Original Series:")
print(s)
print("nResult after replacing all 'ap':")
print(result)

Output result:

原始 Series:
0        apple
1       banana
2       apricot
3    pineapple
dtype: object

替换所有 'ap' 后的结果:
0       XXle
1       banana
2       XXricot
3    XXneXXple

Code explanation:

  • Setregex=Trueto True, then use regular expressions for matching.
  • In a regular expression,'ap'matches all occurrences of 'ap' in the string.
  • 'pineapple' has two occurrences of 'ap', and both are replaced with 'XX'.

Example 3: Replacing Numbers and Symbols

replace()It can be used to clean data, such as removing numbers or special symbols.

Example

import pandas as pd

# Create a Series containing mixed content
s = pd.Series(['phone: 123-456-7890', 'price: $99.99', 'date: 2024-01-01'])

# Remove all numbers
result_digits = s.str.replace(r'd', '', regex=True)

# Remove all non-alphanumeric characters (keep spaces)
result_special = s.str.replace(r'[^a-zA-Z0-9s]', '', regex=True)

print("Original Series:")
print(s)
print("nAfter removing all numbers:")
print(result_digits)
print("nAfter removing all special characters:")
print(result_special)

Output result:

原始 Series:
0    phone: 123-456-7890
1       price: $99.99
2       date: 2024-01-01
dtype: object

移除所有数字后:
phone: ---
price: $.$
date: ------

移除所有特殊字符后:
phone 1234567890
price 9999
date 20240101

Code explanation:

  • r'd'is a regular expression that matches any digit.
  • r'[^a-zA-Z0-9s]'matches all characters that are neither alphanumeric nor spaces.
  • This is a very common operation in data cleaning.

Example 4: Using a Callback Function for Replacement

replThe parameter can be a function for more complex replacement logic.

Example

import pandas as pd

# Create a Series containing text
s = pd.Series(['hello', 'world', 'example', 'python'])

# Use a callback function to convert the matched string to uppercase
result = s.str.replace(r'[aeiou]', lambda m: m.group(0).upper(), regex=True)

print("Original Series:")
print(s)
print("nAfter converting vowels to uppercase:")
print(result)

Output result:

原始 Series:
0       hello
1       world
2      example
3      python
dtype: object

将元音字母转为大写后:
0       hEllo
1       wOrld
2       example
3       pythOn

Code explanation:

  • The callback function receives a match object, and you can usem.group(0)to get the matched content.
  • Converts vowels (a, e, i, o, u) to uppercase.
  • This is a good way to implement more complex replacement logic.

Example 5: Standardizing Data Formats

In real-world projects,replace()it is often used to standardize data formats.

Example

import pandas as pd

# Simulate user-input country names (inconsistent formats)
countries = pd.Series(['USA', 'U.S.A.', 'U.S.A', 'usa', 'us', 'United States'])

# Uniformly replace with standard names
result = countries.replace({
    'USA': 'US',
    'U.S.A.': 'US',
    'U.S.A': 'US',
    'usa': 'US',
    'us': 'US',
    'United States': 'US'
})

print("Original country names:")
print(countries)
print("nResult after unification:")
print(result)

Output result:

原始国家名称:
0           USA
1        U.S.A.
2         U.S.A
3           usa
4            us
5    United States
dtype:
object

统一后的结果:
0    US
1    US
2    US
3    US
4    0    US
dtype: object

Code explanation:

  • This example uses the Seriesreplace()method (not str.replace).
  • You can pass a dictionary to replace multiple values in batch.
  • Unifies various different representations into a standard format.

Notes

  • By default (regex=False),patthe parameter is treated as a plain string, and only the first match is replaced.
  • Setregex=Trueto True, use regular expressions for matching, and all matches can be replaced.
  • Some special characters in regular expressions (such as.、*etc.) need to be escaped.
  • If the Series contains NaN values,replace()it will return NaN.
  • This function returns a new Series and does not modify the original data.

Pandas 常用函数Pandas Common Functions

Other Extensions