Python3 Regular Expressions
A regular expression is a special sequence of characters that helps you conveniently check whether a string matches a certain pattern.
In Python, useremodule to handle regular expressions.
The re module provides a set of functions that allow you to perform pattern matching, search, and replace operations in strings.
reThe module gives the Python language complete regular expression functionality.
This chapter mainly introduces the commonly used regular expression processing functions in Python. If you are not familiar with regular expressions, you can check ourRegular Expression - Tutorial。
re.match function
re.match attempts to match a pattern from the starting position of the string. If the match is not successful at the starting position, match() returns None.
Function syntax:
re.match(pattern, string, flags=0)
Function parameter description:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| flags | Flags, used to control the matching method of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags |
On successful matchre.matchthe method returns a match object, otherwise it returnsNone。
We can usegroup(num)orgroups()match object functions to obtain the matched expression.
| Match object methods | Description |
|---|---|
| group(num=0) | The string of the entire matched expression. group() can input multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups. |
| groups() | Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained. |
Example
The output of the above example is:
(0, 3) None
Example
The execution result of the above example is as follows:
matchObj.group() : Cats are smarter than dogs matchObj.group(1) : Cats matchObj.group(2) : smarter
re.search method
re.search scans the entire string and returns the first successful match.
Function syntax:
re.search(pattern, string, flags=0)
Function parameter description:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| flags | Flags, used to control the matching method of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags |
On successful match, the re.search method returns a match object, otherwise it returns None.
We can use group(num) or groups() match object functions to obtain the matched expression.
| Match object methods | Description |
|---|---|
| group(num=0) | The string of the entire matched expression. group() can input multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups. |
| groups() | Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained. |
Example
The output of the above example is:
(0, 3) (11, 14)
Example
searchObj.group() : Cats are smarter than dogs searchObj.group(1) : Cats searchObj.group(2) : smarter
Difference between re.match and re.search
re.matchIt only matches the beginning of the string. If the beginning of the string does not match the regular expression, the match fails and the function returns None, whilere.searchIt matches the entire string until a match is found.
Example
No match!! search --> matchObj.group() : dogs
Search and Replace
Python's re module provides re.sub for replacing matches in a string.
Syntax:
re.sub(pattern, repl, string, count=0, flags=0)
Parameters:
- pattern: The pattern string in the regular expression.
- repl: The replacement string, which can also be a function.
- string: The original string to be searched and replaced.
- count: The maximum number of replacements after pattern matching. The default 0 means replace all matches.
- flags: The matching mode used at compile time, in numeric form.
The first three are required parameters, and the last two are optional parameters.
Example
The execution result of the above example is as follows:
电话号码 : 2004-959-559 电话号码 : 2004959559
repl parameter is a function
In the following example, the matched numbers in the string are multiplied by 2:
Example
The execution output result is:
A46G8HFD1134
compile function
The compile function is used to compile a regular expression and generate a regular expression (Pattern) object, which is used by the match() and search() functions.
The syntax format is:
re.compile(pattern[, flags])
Parameters:
- pattern: A regular expression in string form
- flags is optional, indicating the matching mode, such as ignoring case, multiline mode, etc. The specific parameters are:
-
re.IGNORECASE or re.I- Makes matching case-insensitive
- re.L indicates that the special character sets \w, \W, \b, \B, \s, \S depend on the current environment
- re.MULTILINE or re.M - Multiline mode, changes the behavior of ^ and $ so that they match the beginning and end of each line of the string.
- re.DOTALL or re.S - Makes.match any character including newline.
- re.ASCII - Makes \w, \W, \b, \B, \d, \D, \s, \S match only ASCII characters.
- re.VERBOSE or re.X - Ignores whitespace and comments, allowing complex regular expressions to be organized more clearly.
These flags can be used individually or combined using bitwise OR (|). For example, re.IGNORECASE | re.MULTILINE means enabling both ignoring case and multiline mode.
Example
Example
Above, when the match succeeds, a Match object is returned, where:
group([group1, …])The method is used to obtain one or more group-matched strings. When you need to obtain the entire matched substring, you can directly usegroup()orgroup(0);start([group])The method is used to obtain the starting position of the group-matched substring in the entire string (the index of the first character of the substring). The default value of the parameter is 0;end([group])The method is used to obtain the ending position of the group-matched substring in the entire string (the index of the last character of the substring + 1). The default value of the parameter is 0;span([group])The method returns(start(group), end(group))。
Let's look at another example:
Example
findall
Find all substrings matched by the regex in the string and return a list. If there are multiple matching patterns, return a list of tuples. If no matches are found, return an empty list.
Note:match and search match once, findall matches all.
The syntax is:
re.findall(pattern, string, flags=0) 或 pattern.findall(string[, pos[, endpos]])
Parameters:
- patternThe matching pattern.
- stringThe string to be matched.
- posOptional parameter, specifies the starting position in the string, defaults to 0.
- endposOptional parameter, specifies the ending position in the string, defaults to the length of the string.
Find all numbers in a string:
Example
Output result:
['123', '456'] ['123', '456'] ['88', '12']
Multiple matching patterns, return a list of tuples:
Example
result = re.findall(r'(\w+)=(\d+)', 'set width=20 and height=10')
print(result)
[('width', '20'), ('height', '10')]
re.finditer
Similar to findall, find all substrings matched by the regex in the string and return them as an iterator.
re.finditer(pattern, string, flags=0)
Parameters:
| Parameter | Description |
|---|---|
| pattern | The regex to match |
| string | The string to be matched. |
| flags | Flags, used to control the matching behavior of the regex, e.g., case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags |
Example
Output result:
12 32 43 3
re.split
The split method splits the string by the matched substrings and returns a list. Its usage is as follows:
re.split(pattern, string[, maxsplit=0, flags=0])
Parameters:
| Parameter | Description |
|---|---|
| pattern | The regex to match |
| string | The string to be matched. |
| maxsplit | Number of splits, maxsplit=1 splits once, default is 0, no limit on the number. |
| flags | Flags, used to control the matching behavior of the regex, e.g., case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags |
Example
Regular Expression Object
re.RegexObject
re.compile() returns a RegexObject object.
re.MatchObject
group() returns the string matched by the RE.
- start()Returns the starting position of the match
- end()Returns the ending position of the match
- span()Returns a tuple containing the (start, end) positions of the match
Regular Expression Modifiers - Optional Flags
Regular expressions can include some optional flag modifiers to control matching modes.
The following flags can be used alone or combined with bitwise OR (|). For example, re.IGNORECASE | re.MULTILINE means enabling both case-insensitive and multiline modes.
| Modifier | Description | Example |
|---|---|---|
| re.IGNORECASE or re.I | Makes matching case-insensitive |
import re
pattern = re.compile(r'apple', flags=re.IGNORECASE)
result = pattern.match('Apple')
print(result.group()) # 输出: 'Apple'
|
| re.MULTILINE or re.M | Multiline matching, affects^and$, making them match the beginning and end of each line in the string. |
import re pattern = re.compile(r'^\d+', flags=re.MULTILINE) text = '123\n456\n789' result = pattern.findall(text) print(result) # 输出: ['123', '456', '789'] |
| re.DOTALL or re.S: | make.Matches any character including newline. |
import re
pattern = re.compile(r'a.b', flags=re.DOTALL)
result = pattern.match('a\nb')
print(result.group()) # 输出: 'a\nb'
|
| re.ASCII | Make \w, \W, \b, \B, \d, \D, \s, \S match only ASCII characters. |
import re
pattern = re.compile(r'\w+', flags=re.ASCII)
result = pattern.match('Hello123')
print(result.group()) # 输出: 'Hello123'
|
| re.VERBOSE or re.X | Ignores whitespace and comments, allowing complex regexes to be organized more clearly. |
import re
pattern = re.compile(r'''
\d+ # 匹配数字
[a-z]+ # 匹配小写字母
''', flags=re.VERBOSE)
result = pattern.match('123abc')
print(result.group()) # 输出: '123abc'
|
Regular Expression Patterns
Pattern strings use special syntax to represent a regular expression.
Letters and numbers represent themselves. Letters and numbers in a regex pattern match the same strings.
Most letters and numbers have different meanings when preceded by a backslash.
Punctuation marks match themselves only when escaped; otherwise they represent special meanings.
The backslash itself needs to be escaped with a backslash.
Since regexes often contain backslashes, it is best to use raw strings to represent them. Pattern elements (such asr'\t', equivalent to\\t) match the corresponding special characters.
The following table lists the special elements in regex pattern syntax. If you provide optional flag parameters while using patterns, the meanings of some pattern elements will change.
| Pattern | Description |
|---|---|
| ^ | Matches the beginning of the string. |
| $ | Matches the end of the string. |
| . | Matches any character except newline. When the re.DOTALL flag is specified, it can match any character including newline. |
| [...] | Used to match any one of the contained characters, e.g., [amk] matches 'a', 'm', or 'k'. |
| [^...] | Characters not in []: [^abc] matches characters other than a, b, c. |
| re* | Matches 0 or more repetitions of the expression. |
| re+ | Matches 1 or more repetitions of the expression. |
| re? | Matches 0 or 1 occurrence of the fragment defined by the preceding regex, non-greedy. |
| re{ n} | Matches n occurrences of the preceding expression. For example, "o{2}" cannot match the "o" in "Bob", but can match the two o's in "food". |
| re{ n,} | Matches n or more of the preceding expression. For example, "o{2,}" cannot match the "o" in "Bob", but can match all the o's in "foooood". "o{1,}" is equivalent to "o+". "o{0,}" is equivalent to "o*". |
| re{ n, m} | Matches n to m repetitions of the fragment defined by the preceding regex, greedy. |
| a| b | Matches a or b |
| (re) | Matches the expression inside the parentheses, and also represents a group. |
| (?imx) | Regex contains three optional flags: i, m, or x. Only affects the area within the parentheses. |
| (?-imx) | Regex turns off the i, m, or x optional flags. Only affects the area within the parentheses. |
| (?: re) | Similar to (...), but does not represent a group. |
| (?imx: re) | Uses i, m, or x optional flags inside the parentheses. |
| (?-imx: re) | Does not use i, m, or x optional flags inside the parentheses. |
| (?#...) | Comment. |
| (?= re) | Positive lookahead assertion. If the contained regex, denoted by ..., matches successfully at the current position, it succeeds; otherwise it fails. However, once the contained expression has been attempted, the matching engine does not advance at all; the rest of the pattern still tries the right side of the assertion. |
| (?! re) | Negative lookahead assertion. The opposite of the positive assertion; succeeds when the contained expression cannot match at the current position in the string. |
| (?> re) | Independent matching pattern, omitting backtracking. |
| \w | Matches digits, letters, and underscore. |
| \W | Matches non-digits, letters, and underscore. |
| \s | Matches any whitespace character, equivalent to [\t\n\r\f]. |
| \S | Matches any non-whitespace character. |
| \d | Matches any digit, equivalent to [0-9]. |
| \D | Matches any non-digit. |
| \A | Matches the beginning of the string. |
| \Z | Matches the end of the string. If there is a newline, it matches only the end of the string before the newline. |
| \z | Matches the end of the string. |
| \G | Matches the position where the last match completed. |
| \b | Matches a word boundary, which is the position between a word and a space. For example, 'er\b' can match the 'er' in "never", but cannot match the 'er' in "verb". |
| \B | Matches a non-word boundary. 'er\B' can match the 'er' in "verb", but cannot match the 'er' in "never". |
| \n, \t, etc. | Matches a newline. Matches a tab, etc. |
| \1...\9 | Matches the content of the nth group. |
| \10 | Matches the content of the nth group if it has matched. Otherwise refers to an octal character code expression. |
Regular Expression Examples
Character matching
| Example | Description |
|---|---|
| python | Matches "python". |
Character classes
| Example | Description |
|---|---|
| [Pp]ython | Matches "Python" or "python" |
| rub[ye] | Matches "ruby" or "rube" |
| [aeiou] | Matches any one letter inside the brackets |
| [0-9] | Matches any digit. Similar to |
| [a-z] | Matches any lowercase letter |
| [A-Z] | Matches any uppercase letter |
| [a-zA-Z0-9] | Matches any letter and digit |
| [^aeiou] | All characters except the letters aeiou |
| [^0-9] | Matches characters other than digits |
Special character classes
| Example | Description |
|---|---|
| . | Matches any single character except "\n". To match any character including '\n', use a pattern like '[.\n]'. |
| \d | Matches a digit character. Equivalent to [0-9]. |
| \D | Matches a non-digit character. Equivalent to [^0-9]. |
| \s | Matches any whitespace character, including space, tab, form feed, etc. Equivalent to [ \f\n\r\t\v]. |
| \S | Matches any non-whitespace character. Equivalent to [^ \f\n\r\t\v]. |
| \w | Matches any word character including underscore. Equivalent to '[A-Za-z0-9_]'. |
| \W | Matches any non-word character. Equivalent to '[^A-Za-z0-9_]'. |