Python3 Regular Expressions

A regular expression is a special sequence of characters that helps you conveniently check whether a string matches a certain pattern.

In Python, useremodule to handle regular expressions.

The re module provides a set of functions that allow you to perform pattern matching, search, and replace operations in strings.

reThe module gives the Python language complete regular expression functionality.

This chapter mainly introduces the commonly used regular expression processing functions in Python. If you are not familiar with regular expressions, you can check ourRegular Expression - Tutorial。


re.match function

re.match attempts to match a pattern from the starting position of the string. If the match is not successful at the starting position, match() returns None.

Function syntax:

re.match(pattern, string, flags=0)

Function parameter description:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
flagsFlags, used to control the matching method of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags

On successful matchre.matchthe method returns a match object, otherwise it returnsNone。

We can usegroup(num)orgroups()match object functions to obtain the matched expression.

Match object methodsDescription
group(num=0)The string of the entire matched expression. group() can input multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups.
groups()Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained.

Example

#!/usr/bin/python import re print(re.match('www', 'www.example.com').span()) # Match at the starting position print(re.match('com', 'www.example.com')) # Not matching at the starting position

The output of the above example is:

(0, 3)
None

Example

#!/usr/bin/python3 import re line = "Cats are smarter than dogs" # .* means matching any single or multiple characters except newline characters (\n, \r) # (.*?) indicates "non-greedy" mode, which only saves the first matched substring matchObj = re.match( r'(.*) are (.*?) .*', line, re.M|re.I) if matchObj: print ("matchObj.group() : ", matchObj.group()) print ("matchObj.group(1) : ", matchObj.group(1)) print ("matchObj.group(2) : ", matchObj.group(2)) else: print ("No match!!")

The execution result of the above example is as follows:

matchObj.group() :  Cats are smarter than dogs
matchObj.group(1) :  Cats
matchObj.group(2) :  smarter

re.search method

re.search scans the entire string and returns the first successful match.

Function syntax:

re.search(pattern, string, flags=0)

Function parameter description:

ParameterDescription
patternThe regular expression to match.
stringThe string to be matched.
flagsFlags, used to control the matching method of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags

On successful match, the re.search method returns a match object, otherwise it returns None.

We can use group(num) or groups() match object functions to obtain the matched expression.

Match object methodsDescription
group(num=0)The string of the entire matched expression. group() can input multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups.
groups()Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained.

Example

#!/usr/bin/python3 import re print(re.search('www', 'www.example.com').span()) # Match at the starting position print(re.search('com', 'www.example.com').span()) # Not matching at the starting position

The output of the above example is:

(0, 3)
(11, 14)

Example

#!/usr/bin/python3 import re line = "Cats are smarter than dogs" searchObj = re.search( r'(.*) are (.*?) .*', line, re.M|re.I) if searchObj: print ("searchObj.group() : ", searchObj.group()) print ("searchObj.group(1) : ", searchObj.group(1)) print ("searchObj.group(2) : ", searchObj.group(2)) else: print ("Nothing found!!")
The execution result of the above example is as follows:
searchObj.group() :  Cats are smarter than dogs
searchObj.group(1) :  Cats
searchObj.group(2) :  smarter

Difference between re.match and re.search

re.matchIt only matches the beginning of the string. If the beginning of the string does not match the regular expression, the match fails and the function returns None, whilere.searchIt matches the entire string until a match is found.

Example

#!/usr/bin/python3 import re line = "Cats are smarter than dogs" matchObj = re.match( r'dogs', line, re.M|re.I) if matchObj: print ("match --> matchObj.group() : ", matchObj.group()) else: print ("No match!!") matchObj = re.search( r'dogs', line, re.M|re.I) if matchObj: print ("search --> matchObj.group() : ", matchObj.group()) else: print ("No match!!")
The output of the above example is:
No match!!
search --> matchObj.group() :  dogs

Search and Replace

Python's re module provides re.sub for replacing matches in a string.

Syntax:

re.sub(pattern, repl, string, count=0, flags=0)

Parameters:

  • pattern: The pattern string in the regular expression.
  • repl: The replacement string, which can also be a function.
  • string: The original string to be searched and replaced.
  • count: The maximum number of replacements after pattern matching. The default 0 means replace all matches.
  • flags: The matching mode used at compile time, in numeric form.

The first three are required parameters, and the last two are optional parameters.

Example

#!/usr/bin/python3 import re phone = "2004-959-559 # This is a phone number" # Delete comments num = re.sub(r'#.*$', "", phone) print ("Phone number:", num) # Remove non-digit content num = re.sub(r'\D', "", phone) print ("Phone number:", num)

The execution result of the above example is as follows:

电话号码 :  2004-959-559 
电话号码 :  2004959559

repl parameter is a function

In the following example, the matched numbers in the string are multiplied by 2:

Example

#!/usr/bin/python import re # Multiply the matched numbers by 2 def double(matched): value = int(matched.group('value')) return str(value * 2) s = 'A23G4HFD567' print(re.sub('(?P<value>\d+)', double, s))

The execution output result is:

A46G8HFD1134

compile function

The compile function is used to compile a regular expression and generate a regular expression (Pattern) object, which is used by the match() and search() functions.

The syntax format is:

re.compile(pattern[, flags])

Parameters:

  • pattern: A regular expression in string form
  • flags is optional, indicating the matching mode, such as ignoring case, multiline mode, etc. The specific parameters are:
    • re.IGNORECASE or re.I- Makes matching case-insensitive
  • re.L indicates that the special character sets \w, \W, \b, \B, \s, \S depend on the current environment
  • re.MULTILINE or re.M - Multiline mode, changes the behavior of ^ and $ so that they match the beginning and end of each line of the string.
  • re.DOTALL or re.S - Makes.match any character including newline.
  • re.ASCII - Makes \w, \W, \b, \B, \d, \D, \s, \S match only ASCII characters.
  • re.VERBOSE or re.X - Ignores whitespace and comments, allowing complex regular expressions to be organized more clearly.

These flags can be used individually or combined using bitwise OR (|). For example, re.IGNORECASE | re.MULTILINE means enabling both ignoring case and multiline mode.

Example

Example

>>>import re >>> pattern = re.compile(r'\d+') # Used to match at least one digit >>> m = pattern.match('one12twothree34four') # Search at the beginning, no match >>> print( m ) None >>> m = pattern.match('one12twothree34four', 2, 10) # Start matching from the position of 'e', no match >>> print( m ) None >>> m = pattern.match('one12twothree34four', 3, 10) # Start matching from the position of '1', exactly matches >>> print( m ) # Returns a Match object <_sre.SRE_Match object at 0x10a42aac0> >>> m.group(0) # 0 can be omitted '12' >>> m.start(0) # 0 can be omitted 3 >>> m.end(0) # 0 can be omitted 5 >>> m.span(0) # 0 can be omitted (3, 5)

Above, when the match succeeds, a Match object is returned, where:

  • group([group1, …])The method is used to obtain one or more group-matched strings. When you need to obtain the entire matched substring, you can directly usegroup()orgroup(0);
  • start([group])The method is used to obtain the starting position of the group-matched substring in the entire string (the index of the first character of the substring). The default value of the parameter is 0;
  • end([group])The method is used to obtain the ending position of the group-matched substring in the entire string (the index of the last character of the substring + 1). The default value of the parameter is 0;
  • span([group])The method returns(start(group), end(group))。

Let's look at another example:

Example

>>>import re >>> pattern = re.compile(r'([a-z]+) ([a-z]+)', re.I) # re.I means ignoring case >>> m = pattern.match('Hello World Wide Web') >>> print( m ) # Match successful, returns a Match object <_sre.SRE_Match object at 0x10bea83e8> >>> m.group(0) # Return the entire substring that matched successfully 'Hello World' >>> m.span(0) # Return the index of the entire matched substring (0, 11) >>> m.group(1) # Return the substring matched by the first group 'Hello' >>> m.span(1) # Return the index of the substring matched by the first group (0, 5) >>> m.group(2) # Return the substring matched by the second group 'World' >>> m.span(2) # Return the index of the substring matched by the second group (6, 11) >>> m.groups() # Equivalent to (m.group(1), m.group(2), ...) ('Hello', 'World') >>> m.group(3) # The third group does not exist Traceback (most recent call last): File "<stdin>", line 1, in <module> IndexError: no such group

findall

Find all substrings matched by the regex in the string and return a list. If there are multiple matching patterns, return a list of tuples. If no matches are found, return an empty list.

Note:match and search match once, findall matches all.

The syntax is:

re.findall(pattern, string, flags=0)
或
pattern.findall(string[, pos[, endpos]])

Parameters:

  • patternThe matching pattern.
  • stringThe string to be matched.
  • posOptional parameter, specifies the starting position in the string, defaults to 0.
  • endposOptional parameter, specifies the ending position in the string, defaults to the length of the string.

Find all numbers in a string:

Example

import re result1 = re.findall(r'\d+','example 123 google 456') pattern = re.compile(r'\d+') # Find numbers result2 = pattern.findall('example 123 google 456') result3 = pattern.findall('run88oob123google456', 0, 10) print(result1) print(result2) print(result3)

Output result:

['123', '456']
['123', '456']
['88', '12']

Multiple matching patterns, return a list of tuples:

Example

import re

result = re.findall(r'(\w+)=(\d+)', 'set width=20 and height=10')
print(result)
[('width', '20'), ('height', '10')]

re.finditer

Similar to findall, find all substrings matched by the regex in the string and return them as an iterator.

re.finditer(pattern, string, flags=0)

Parameters:

ParameterDescription
patternThe regex to match
stringThe string to be matched.
flagsFlags, used to control the matching behavior of the regex, e.g., case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags

Example

import re it = re.finditer(r"\d+","12a32bc43jf3") for match in it: print (match.group() )

Output result:

12 
32 
43 
3

re.split

The split method splits the string by the matched substrings and returns a list. Its usage is as follows:

re.split(pattern, string[, maxsplit=0, flags=0])

Parameters:

ParameterDescription
patternThe regex to match
stringThe string to be matched.
maxsplitNumber of splits, maxsplit=1 splits once, default is 0, no limit on the number.
flagsFlags, used to control the matching behavior of the regex, e.g., case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags

Example

>>>import re >>> re.split('\W+', 'example, example, example.') ['example', 'example', 'example', ''] >>> re.split('(\W+)', ' example, example, example.') ['', ' ', 'example', ', ', 'example', ', ', 'example', '.', ''] >>> re.split('\W+', ' example, example, example.', 1) ['', 'example, example, example.'] >>> re.split('a*', 'hello world') # For a string that cannot find a match, split will not split it ['hello world']

Regular Expression Object

re.RegexObject

re.compile() returns a RegexObject object.

re.MatchObject

group() returns the string matched by the RE.

  • start()Returns the starting position of the match
  • end()Returns the ending position of the match
  • span()Returns a tuple containing the (start, end) positions of the match

Regular Expression Modifiers - Optional Flags

Regular expressions can include some optional flag modifiers to control matching modes.

The following flags can be used alone or combined with bitwise OR (|). For example, re.IGNORECASE | re.MULTILINE means enabling both case-insensitive and multiline modes.

ModifierDescriptionExample
re.IGNORECASE or re.IMakes matching case-insensitive
import re
pattern = re.compile(r'apple', flags=re.IGNORECASE)
result = pattern.match('Apple')
print(result.group())  # 输出: 'Apple'
re.MULTILINE or re.MMultiline matching, affects^and$, making them match the beginning and end of each line in the string.
import re
pattern = re.compile(r'^\d+', flags=re.MULTILINE)
text = '123\n456\n789'
result = pattern.findall(text)
print(result)  # 输出: ['123', '456', '789']
re.DOTALL or re.S:make.Matches any character including newline.
import re
pattern = re.compile(r'a.b', flags=re.DOTALL)
result = pattern.match('a\nb')
print(result.group())  # 输出: 'a\nb'
re.ASCIIMake \w, \W, \b, \B, \d, \D, \s, \S match only ASCII characters.
import re
pattern = re.compile(r'\w+', flags=re.ASCII)
result = pattern.match('Hello123')
print(result.group())  # 输出: 'Hello123'
re.VERBOSE or re.XIgnores whitespace and comments, allowing complex regexes to be organized more clearly.
import re
pattern = re.compile(r'''
    \d+  # 匹配数字
    [a-z]+  # 匹配小写字母
''', flags=re.VERBOSE)
result = pattern.match('123abc')
print(result.group())  # 输出: '123abc'

Regular Expression Patterns

Pattern strings use special syntax to represent a regular expression.

Letters and numbers represent themselves. Letters and numbers in a regex pattern match the same strings.

Most letters and numbers have different meanings when preceded by a backslash.

Punctuation marks match themselves only when escaped; otherwise they represent special meanings.

The backslash itself needs to be escaped with a backslash.

Since regexes often contain backslashes, it is best to use raw strings to represent them. Pattern elements (such asr'\t', equivalent to\\t) match the corresponding special characters.

The following table lists the special elements in regex pattern syntax. If you provide optional flag parameters while using patterns, the meanings of some pattern elements will change.

PatternDescription
^Matches the beginning of the string.
$Matches the end of the string.
.Matches any character except newline. When the re.DOTALL flag is specified, it can match any character including newline.
[...]Used to match any one of the contained characters, e.g., [amk] matches 'a', 'm', or 'k'.
[^...]Characters not in []: [^abc] matches characters other than a, b, c.
re*Matches 0 or more repetitions of the expression.
re+Matches 1 or more repetitions of the expression.
re?Matches 0 or 1 occurrence of the fragment defined by the preceding regex, non-greedy.
re{ n}Matches n occurrences of the preceding expression. For example, "o{2}" cannot match the "o" in "Bob", but can match the two o's in "food".
re{ n,}Matches n or more of the preceding expression. For example, "o{2,}" cannot match the "o" in "Bob", but can match all the o's in "foooood". "o{1,}" is equivalent to "o+". "o{0,}" is equivalent to "o*".
re{ n, m}Matches n to m repetitions of the fragment defined by the preceding regex, greedy.
a| bMatches a or b
(re)Matches the expression inside the parentheses, and also represents a group.
(?imx)Regex contains three optional flags: i, m, or x. Only affects the area within the parentheses.
(?-imx)Regex turns off the i, m, or x optional flags. Only affects the area within the parentheses.
(?: re)Similar to (...), but does not represent a group.
(?imx: re)Uses i, m, or x optional flags inside the parentheses.
(?-imx: re)Does not use i, m, or x optional flags inside the parentheses.
(?#...)Comment.
(?= re)Positive lookahead assertion. If the contained regex, denoted by ..., matches successfully at the current position, it succeeds; otherwise it fails. However, once the contained expression has been attempted, the matching engine does not advance at all; the rest of the pattern still tries the right side of the assertion.
(?! re)Negative lookahead assertion. The opposite of the positive assertion; succeeds when the contained expression cannot match at the current position in the string.
(?> re)Independent matching pattern, omitting backtracking.
\wMatches digits, letters, and underscore.
\WMatches non-digits, letters, and underscore.
\sMatches any whitespace character, equivalent to [\t\n\r\f].
\SMatches any non-whitespace character.
\dMatches any digit, equivalent to [0-9].
\DMatches any non-digit.
\AMatches the beginning of the string.
\ZMatches the end of the string. If there is a newline, it matches only the end of the string before the newline.
\zMatches the end of the string.
\GMatches the position where the last match completed.
\bMatches a word boundary, which is the position between a word and a space. For example, 'er\b' can match the 'er' in "never", but cannot match the 'er' in "verb".
\BMatches a non-word boundary. 'er\B' can match the 'er' in "verb", but cannot match the 'er' in "never".
\n, \t, etc.Matches a newline. Matches a tab, etc.
\1...\9Matches the content of the nth group.
\10Matches the content of the nth group if it has matched. Otherwise refers to an octal character code expression.

Regular Expression Examples

Character matching

ExampleDescription
pythonMatches "python".

Character classes

ExampleDescription
[Pp]ython Matches "Python" or "python"
rub[ye]Matches "ruby" or "rube"
[aeiou]Matches any one letter inside the brackets
[0-9]Matches any digit. Similar to
[a-z]Matches any lowercase letter
[A-Z]Matches any uppercase letter
[a-zA-Z0-9]Matches any letter and digit
[^aeiou]All characters except the letters aeiou
[^0-9]Matches characters other than digits

Special character classes

ExampleDescription
.Matches any single character except "\n". To match any character including '\n', use a pattern like '[.\n]'.
\dMatches a digit character. Equivalent to [0-9].
\D Matches a non-digit character. Equivalent to [^0-9].
\sMatches any whitespace character, including space, tab, form feed, etc. Equivalent to [ \f\n\r\t\v].
\S Matches any non-whitespace character. Equivalent to [^ \f\n\r\t\v].
\wMatches any word character including underscore. Equivalent to '[A-Za-z0-9_]'.
\WMatches any non-word character. Equivalent to '[^A-Za-z0-9_]'.
Other extensions