Python re Module

Python'sremodule is a standard library module for handling regular expressions.

Regular expressions (Regular Expression, abbreviated as regex or regexp) are a powerful tool for matching, searching, and manipulating text.

By usingrethe module, you can use regular expressions in Python to process strings.

Why use the re module?

In text processing, we often need to find specific patterns or replace certain characters. For example, validating email addresses, extracting links from web pages, or formatting text. Writing code manually to accomplish these tasks can be very tedious, while regular expressions provide a concise and efficient way to solve these problems.


Basic usage of the re module

1. Import the re module

Before usingrethe module, you first need to import it:

import re

2. Common re module functions

2.1 re.match()

re.match()The function is used to match a regular expression from the beginning of a string. If the match succeeds, it returns a match object; otherwise, it returnsNone。

Example

import re

pattern = r"hello"
text = "hello world"

match = re.match(pattern, text)
if match:
    print("Match successful:", match.group())
else:
    print("Match failed")

Output:

匹配成功: hello

2.2 re.search()

re.search()The function is used to search for the first match of a regular expression in a string. Unlikere.match()different,re.search()does not require the match to start from the beginning of the string.

Example

import re

pattern = r"world"
text = "hello world"

match = re.search(pattern, text)
if match:
    print("Match successful:", match.group())
else:
    print("Match failed")

Output:

匹配成功: world

2.3 re.findall()

re.findall()The function is used to find all substrings in a string that match the regular expression, and returns a list.

Example

import re

pattern = r"\d+"
text = "There are 3 apples and 5 oranges."

matches = re.findall(pattern, text)
print("Numbers found:", matches)

Output:

找到的数字: ['3', '5']

2.4 re.sub()

re.sub()The function is used to replace the parts of a string that match the regular expression.

Example

import re

pattern = r"apple"
text = "I have an apple."

new_text = re.sub(pattern, "banana", text)
print("Replaced text:", new_text)

Output:

替换后的文本: I have an banana.

Basic syntax of regular expressions

1. Ordinary characters

Ordinary characters (such as letters, digits) match themselves directly.

Example

import re

pattern = r"cat"
text = "The cat is on the mat."

match = re.search(pattern, text)
if match:
    print("Match successful:", match.group())

Output:

匹配成功: cat

2. Special characters

There are some special characters in regular expressions that have special meanings. For example:

  • .: Matches any single character (except newline).
  • *: Matches the preceding character zero or more times.
  • +: Matches the preceding character one or more times.
  • ?: Matches the preceding character zero or one time.
  • \d: Matches any digit character (equivalent to[0-9])。
  • \w: Matches any letter, digit, or underscore character (equivalent to[a-zA-Z0-9_])。

Example

import re

pattern = r"\d+"
text = "The price is 100 dollars."

match = re.search(pattern, text)
if match:
    print("Match successful:", match.group())

Output:

匹配成功: 100

3. Character sets

Character sets are used to match any one character from a set. For example,[abc]matchesa、borc。

Example

import re

pattern = r"[aeiou]"
text = "Hello World!"

matches = re.findall(pattern, text)
print("Vowels found:", matches)

Output:

找到的元音字母: ['e', 'o', 'o']

4. Grouping

Grouping allows you to combine multiple characters together and operate on them. For example,(abc)matchesabc。

Example

import re

pattern = r"(ab)+"
text = "ababab"

match = re.search(pattern, text)
if match:
    print("Match successful:", match.group())

Output:

匹配成功: ababab

Practical exercises

Exercise 1: Validate email addresses

Write a regular expression to validate the format of an email address.

Example

import re

pattern = r"^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+$"
email = "example@example.com"

if re.match(pattern, email):
    print("Valid email address")
else:
    print("Invalid email address")

Exercise 2: Extract phone numbers

Write a regular expression to extract phone numbers from text.

Example

import re

pattern = r"\d{3}-\d{3}-\d{4}"
text = "My phone number is 123-456-7890."

match = re.search(pattern, text)
if match:
    print("Phone numbers found:", match.group())

1. Core functions

Method Description Example
re.compile(pattern) Precompile a regular expression (improves reuse performance) pat = re.compile(r'\d+')
re.search(pattern, string) Search for the first match in a string re.search(r'\d+', 'a1b2')→ matches'1'
re.match(pattern, string) Match from the beginning of the string re.match(r'\d+', '123a')→ matches'123'
re.fullmatch(pattern, string) The entire string fully matches re.fullmatch(r'\d+', '123')→ matches'123'
re.findall(pattern, string) Return a list of all non-overlapping matches re.findall(r'\d+', 'a1b22c') → ['1', '22']
re.finditer(pattern, string) Return an iterator of all matches (including position information) for m in re.finditer(r'\d+', 'a1b2'): print(m.group())
re.sub(pattern, repl, string) Replace matches re.sub(r'\d+', 'X', 'a1b2') → 'aXbX'
re.split(pattern, string) Split the string by matches re.split(r'\d+', 'a1b2c') → ['a', 'b', 'c']
re.escape() Escape special characters re.escape("C:\\Users\\test.txt")
re.purge() Clear cache re.purge()

2. Match object (Match) methods/attributes

Method/Attribute Description Example
group() Return the entire matched string m.group() → 'abc'
group(n) Return the content of the nth capture group m = re.search(r'(\d)(\d)', '12'); m.group(1) → '1'
groups() Return a tuple of all capture groups m.groups() → ('1', '2')
start()/end() Start/end position of the match m.start() → 0
span() Return the match range(start, end) m.span() → (0, 2)

3. Regular expression metacharacters (partial)

Metacharacter Description Example match
. Match any character (except newline) a.c → 'abc'
\d Match a digit \d+ → '123'
\D Match a non-digit \D+ → 'abc'
\w Match a word character (letter, digit, underscore) \w+ → 'Ab_1'
\W Match a non-word character \W+ → '!@#'
\s Match a whitespace character (space, tab, etc.) \s+ → ' \t'
\S Match a non-whitespace character \S+ → 'abc'
[] Character set [A-Za-z]→ any letter
^ Match the beginning of a string ^\d+→ digit at the beginning
$ Match the end of a string \d+$→ digit at the end
* Match the preceding character zero or more times a* → '', 'aaa'
+ Match the preceding character one or more times a+ → 'a', 'aaa'
? Match the preceding character zero or one time a? → '', 'a'
{m,n} Match the preceding character m to n times a{2,3} → 'aa', 'aaa'
| OR operation cat|dog → 'cat'or'dog'
() Capture group (\d+)→ extract digits

4. Compilation flags (flagsparameter)

Flag Description Example
re.IGNORECASE (re.I) Ignore case re.search(r'abc', 'ABC', re.I)
re.MULTILINE (re.M) Multiline mode (affects^and$) re.findall(r'^\d+', '1\n2', re.M) → ['1', '2']
re.DOTALL (re.S) Let.Match all characters including newlines re.search(r'a.*b', 'a\nb', re.S)
re.ASCII Let\w, \Wetc. only match ASCII characters re.search(r'\w+', 'こん', re.ASCII)→ no match
re.VERBOSE (re.X) Allow comments and spaces in the regex re.compile(r'''\d+ # 匹配数字''', re.X)

Example

1. Extract email addresses

Example

import re

text = "Contact: admin@example.com, support@test.org"
emails = re.findall(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b', text)
print(emails)  # Output: ['admin@example.com', 'support@test.org']

2. Replace date formats

Example

date_str = "Today is 05-15-2023"
new_str = re.sub(r'(\d{2})-(\d{2})-(\d{4})', r'\3year\1month\2day', date_str)
print(new_str)  # Output: "Today is 2023-05-15"

3. Multi-condition matching

Example

pattern = re.compile(r'''
    ^(?P<username>\w+) # username
    :(?P<password>\S+) # password
    @(?P<domain>\w+\.\w+) # domain
$'''
, re.VERBOSE)

m = pattern.match("john:pass123@example.com")
if m:
    print(m.groupdict())  # Output: {'username': 'john', 'password': 'pass123', 'domain': 'example.com'}

4. Split complex strings

Example

text = "Apple1Banana2Cherry3Date"
parts = re.split(r'\d+', text)
print(parts)  # Output: ['Apple', 'Banana', 'Cherry', 'Date']

Notes

  1. Raw strings: It is recommended to use raw strings for regular expressions (r'\d'), to avoid escape character conflicts.

  2. Greedy matching: Greedy matching by default (e.g..*will match the longest possible), you need to use.*?to implement non-greedy matching.

  3. Performance Optimization: Frequently used regular expressions should prioritize usingre.compile()precompilation.

  4. Backtracking Issues: Complex regular expressions may cause performance issues (such as nested quantifiers(a+)+)。

Other Extensions