Python Regular Expressions
A regular expression is a special sequence of characters that helps you easily check whether a string matches a certain pattern.
Python added the re module since version 1.5, which provides Perl-style regular expression patterns.
The re module gives the Python language all the functionality of regular expressions.
The compile function generates a regular expression object based on a pattern string and optional flag parameters. This object has a series of methods for regular expression matching and replacement.
The re module also provides functions that are fully consistent with the functionality of these methods. These functions use a pattern string as their first parameter.
This chapter mainly introduces the commonly used regular expression processing functions in Python.
re.match function
re.match attempts to match a pattern from the start of the string. If the match is not successful at the starting position, match() returns None.
Function syntax:
re.match(pattern, string, flags=0)
Function parameter description:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| flags | Flags, used to control the matching method of the regular expression, such as: case sensitivity, multiline matching, etc. See:Regular Expression Modifiers - Optional Flags |
If the match is successful, the re.match method returns a match object, otherwise it returns None.
We can use group(num) or groups() match object functions to obtain the matched expression.
| Match object methods | Description |
|---|---|
| group(num=0) | The string of the entire matched expression. group() can take multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups. |
| groups() | Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained. |
Example
The output of the above example is:
(0, 3) None
Example
The result of the above example is as follows:
matchObj.group() : Cats are smarter than dogs matchObj.group(1) : Cats matchObj.group(2) : smarter
re.search method
re.search scans the entire string and returns the first successful match.
Function syntax:
re.search(pattern, string, flags=0)
Function parameter description:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| flags | Flags, used to control the matching method of the regular expression, such as: case sensitivity, multiline matching, etc. |
If the match is successful, the re.search method returns a match object, otherwise it returns None.
We can use group(num) or groups() match object functions to obtain the matched expression.
| Match object methods | Description |
|---|---|
| group(num=0) | The string of the entire matched expression. group() can take multiple group numbers at once, in which case it returns a tuple containing the values corresponding to those groups. |
| groups() | Returns a tuple containing all subgroup strings, from 1 to the number of subgroups contained. |
Example
The output of the above example is:
(0, 3) (11, 14)
Example
searchObj.group() : Cats are smarter than dogs searchObj.group(1) : Cats searchObj.group(2) : smarter
Difference between re.match and re.search
re.match only matches the beginning of the string. If the beginning of the string does not conform to the regular expression, the match fails and the function returns None; while re.search matches the entire string until a match is found.
Example
No match!! search --> searchObj.group() : dogs
Search and Replace
Python's re module provides re.sub for replacing matches in a string.
Syntax:
re.sub(pattern, repl, string, count=0, flags=0)
Parameters:
- pattern: The pattern string in the regular expression.
- repl: The replacement string, can also be a function.
- string: The original string to be searched and replaced.
- count: The maximum number of replacements after pattern matching, default 0 means replace all matches.
Example
电话号码是: 2004-959-559 电话号码是 : 2004959559
repl parameter is a function
In the following example, the matched numbers in the string are multiplied by 2:
Example
The output after execution is:
A46G8HFD1134
re.compile function
The compile function is used to compile a regular expression and generate a regular expression (Pattern) object, which can be used by functions such as match(), search(), and findall.
The syntax format is:
re.compile(pattern[, flags])
Parameters:
-
pattern: A regular expression in string form
-
flags: Optional, representing the matching mode, such as ignoring case, multiline mode, etc. The specific parameters are:
- re.IIgnore case
- re.LIndicates that the special character sets \w, \W, \b, \B, \s, \S depend on the current environment
- re.MMultiline mode
- re.SThat is.And any character including newline (.Excluding newline)
- re.UIndicates that the special character sets \w, \W, \b, \B, \d, \D, \s, \S depend on the Unicode character property database
- re.XTo increase readability, ignore spaces and#Comments after
Example
Example
Above, upon a successful match, a Match object is returned, where:
group([group1, …])The method is used to obtain the string matched by one or more groups. When you want to obtain the entire matched substring, you can directly usegroup()orgroup(0);start([group])The method is used to get the starting position of the substring matched by a group in the entire string (the index of the first character of the substring), with a default parameter value of 0;end([group])The method is used to get the ending position of the substring matched by a group in the entire string (the index of the last character of the substring + 1), with a default parameter value of 0;span([group])The method returns(start(group), end(group))。
Let's look at another example:
Example
findall
Finds all substrings matched by the regular expression in the string and returns a list. If there are multiple matching patterns, it returns a list of tuples. If no matches are found, it returns an empty list.
Note:match and search match once, while findall matches all.
The syntax format is:
findall(string[, pos[, endpos]])
Parameters:
- string: The string to be matched.
- pos: Optional parameter, specifies the starting position of the string, defaults to 0.
- endpos: Optional parameter, specifies the ending position of the string, defaults to the length of the string.
Find all numbers in the string:
Example
Output result:
['123', '456'] ['88', '12']
Multiple matching patterns, returns a list of tuples:
Example
result = re.findall(r'(\w+)=(\d+)', 'set width=20 and height=10')
print(result)
[('width', '20'), ('height', '10')]
re.finditer
Similar to findall, finds all substrings matched by the regular expression in the string and returns them as an iterator.
re.finditer(pattern, string, flags=0)
Parameters:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| flags | Flags, used to control the matching behavior of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags |
Example
Output result:
12 32 43 3
re.split
The split method splits the string according to the substrings that can be matched and returns a list. Its usage form is as follows:
re.split(pattern, string[, maxsplit=0, flags=0])
Parameters:
| Parameter | Description |
|---|---|
| pattern | The regular expression to match. |
| string | The string to be matched. |
| maxsplit | Maximum number of splits. maxsplit=1 means split once. Default is 0, which means no limit. |
| flags | Flags, used to control the matching behavior of the regular expression, such as case sensitivity, multiline matching, etc. See:Regular expression modifiers - optional flags |
Example
Regular Expression Object
re.RegexObject
re.compile() returns a RegexObject object.
re.MatchObject
group() returns the string matched by the RE.
- start()Returns the starting position of the match
- end()Returns the ending position of the match
- span()Returns a tuple containing the (start, end) positions of the match
Regular Expression Modifiers - Optional Flags
Regular expressions can contain optional flag modifiers to control the matching pattern. Modifiers are specified as an optional flag. Multiple flags can be specified by bitwise ORing (|) them. For example, re.I | re.M sets both the I and M flags:
| Modifier | Description |
|---|---|
| re.I | Makes the matching case-insensitive |
| re.L | Performs locale-aware matching |
| re.M | Multi-line matching, affects ^ and $ |
| re.S | Makes . match all characters including newline |
| re.U | Parses characters according to the Unicode character set. This flag affects \w, \W, \b, \B. |
| re.X | This flag gives you a more flexible format so you can write regular expressions that are easier to understand. |
Regular Expression Patterns
Pattern strings use special syntax to represent a regular expression:
Letters and numbers represent themselves. Letters and numbers in a regular expression pattern match the same strings.
Most letters and numbers have different meanings when preceded by a backslash.
Punctuation characters match themselves only when escaped; otherwise they represent special meanings.
The backslash itself needs to be escaped with a backslash.
Since regular expressions often contain backslashes, you'd better use raw strings to represent them. Pattern elements (such as r'\t', equivalent to '\\t') match the corresponding special characters.
The following table lists the special elements in regular expression pattern syntax. If you provide optional flag parameters while using a pattern, the meaning of certain pattern elements may change.
| Pattern | Description |
|---|---|
| ^ | Matches the beginning of the string. |
| $ | Matches the end of the string. |
| . | Matches any character except newline. When the re.DOTALL flag is specified, it can match any character including newline. |
| [...] | Used to represent a set of characters, listed individually: [amk] matches 'a', 'm', or 'k' |
| [^...] | Characters not in []: [^abc] matches any character except a, b, c. |
| re* | Matches 0 or more occurrences of the expression. |
| re+ | Matches 1 or more occurrences of the expression. |
| re? | Matches 0 or 1 occurrences of the fragment defined by the preceding regular expression, non-greedy way |
| re{ n} | Exactly matches n occurrences of the preceding expression. For example,o{2}It cannot match the 'o' in "Bob", but it can match the two o's in "food". |
| re{ n,} | Matches n or more occurrences of the preceding expression. For example, o{2,} cannot match the 'o' in "Bob", but it can match all the o's in "foooood". "o{1,}" is equivalent to "o+". "o{0,}" is equivalent to "o*". |
| re{ n, m} | Matches n to m occurrences of the fragment defined by the preceding regular expression, greedy way |
| a| b | Matches a or b |
| (re) | Groups the regular expression and remembers the matched text |
| (?imx) | Regular expression contains three optional flags: i, m, or x. Only affects the area within the parentheses. |
| (?-imx) | Regular expression turns off i, m, or x optional flags. Only affects the area within the parentheses. |
| (?: re) | Similar to (...), but does not represent a group |
| (?imx: re) | Uses i, m, or x optional flags inside the parentheses |
| (?-imx: re) | Does not use i, m, or x optional flags inside the parentheses |
| (?#...) | Comment. |
| (?= re) | Positive lookahead assertion. If the contained regular expression, represented by ..., matches successfully at the current position, it succeeds; otherwise it fails. But once the contained expression has been attempted, the matching engine does not advance; the rest of the pattern still has to try to the right of the assertion. |
| (?! re) | Negative lookahead assertion. Opposite of the positive assertion; succeeds when the contained expression cannot match at the current position in the string. |
| (?> re) | Matches an independent pattern, eliminating backtracking. |
| \w | Matches letters, digits, and underscores |
| \W | Matches anything except letters, digits, and underscores |
| \s | Matches any whitespace character, equivalent to[ \t\n\r\f]。 |
| \S | Matches any non-whitespace character |
| \d | Matches any digit, equivalent to [0-9]. |
| \D | Matches any non-digit |
| \A | Matches the start of the string |
| \Z | Matches the end of the string. If there is a newline, it only matches up to the end of the string before the newline. |
| \z | Matches the end of the string |
| \G | Matches the position where the last match completed. |
| \b | Matches a word boundary, that is, the position between a word and a space. For example, 'er\b' can match the 'er' in "never", but cannot match the 'er' in "verb". |
| \B | Matches a non-word boundary. 'er\B' can match the 'er' in "verb", but cannot match the 'er' in "never". |
| \n, \t, etc. | Matches a newline character. Matches a tab character. etc. |
| \1...\9 | Matches the content of the nth group. |
| \10 | Matches the content of the nth group if it has been matched. Otherwise it refers to an octal character code expression. |
Regular Expression Examples
Character matching
| Example | Description |
|---|---|
| python | Matches "python". |
Character classes
| Example | Description |
|---|---|
| [Pp]ython | Matches "Python" or "python" |
| rub[ye] | Matches "ruby" or "rube" |
| [aeiou] | Matches any one letter in the brackets |
| [0-9] | Matches any digit. Similar to |
| [a-z] | Matches any lowercase letter |
| [A-Z] | Matches any uppercase letter |
| [a-zA-Z0-9] | Matches any letter and digit |
| [^aeiou] | All characters except the letters aeiou |
| [^0-9] | Matches characters other than digits |
Special character classes
| Example | Description |
|---|---|
| . | Matches any single character except "\n". To match any character including '\n', use a pattern like '[.\n]'. |
| \d | Matches a digit character. Equivalent to [0-9]. |
| \D | Matches a non-digit character. Equivalent to [^0-9]. |
| \s | Matches any whitespace character, including space, tab, form feed, etc. Equivalent to [ \f\n\r\t\v]. |
| \S | Matches any non-whitespace character. Equivalent to [^ \f\n\r\t\v]. |
| \w | Matches any word character including underscore. Equivalent to '[A-Za-z0-9_]'. |
| \W | Matches any non-word character. Equivalent to '[^A-Za-z0-9_]'. |