Ruby Regular Expressions

Regular Expressionsis a special sequence of characters that matches or searches a collection of strings by using a pattern with specialized syntax.

Regular expressions use a predefined set of specific characters and combinations of these characters to form a "rule string", which is used to express a filtering logic for strings.

Syntax

Regular ExpressionsLiterally, it is a pattern between slashes or between arbitrary delimiters following %r, as shown below:

/pattern/ /pattern/im #Options can be specified %r!/usr/local! #Regular Expressions with Delimiters

Examples

#!/usr/bin/ruby line1 = "Cats are smarter than dogs"; line2 = "Dogs also like meat"; if ( line1 =~ /Cats(.*)/ ) puts "Line1 contains Cats" end if ( line2 =~ /Cats(.*)/ ) puts "Line2 contains Dogs" end

Try it »

The output of the above example is:

Line1 contains Cats

Regular Expression Modifiers

A regular expression literally may contain an optional modifier to control various aspects of matching. The modifier is specified after the second slash character, as shown in the example above. The following table lists the possible modifiers:

ModifierDescription
iIgnore case when matching text.
oPerform #{} interpolation only once; the regular expression is evaluated the first time.
xIgnore spaces, allowing whitespace and comments to be placed within the entire expression.
mMatch multiple lines, recognizing newline characters as normal characters.
u,e,s,nInterpret the regular expression as Unicode (UTF-8), EUC, SJIS, or ASCII. If no modifier is specified, the regular expression is considered to use the source encoding.

Just as strings are delimited by %Q, Ruby allows you to use %r as the start of a regular expression, followed by any delimiter. This is very useful when describing patterns that contain many slash characters you don't want to escape.

#The following matches a single slash character without escaping: %r|/| # Flag characters can be matched using the following syntax %r[</(.*)>]i

Regular Expression Patterns

Except for control characters,(+ ? . * ^ $ ( ) [ ] { } | \)all other characters match themselves. You can escape control characters by placing a backslash before them.

The following table lists the regular expression syntax available in Ruby.

PatternDescription
^Matches the beginning of a line.
$Matches the end of a line.
.Matches any single character except newline. With the m option, it also matches newline.
[...]Matches any single character in square brackets.
[^...]Matches any single character not in square brackets.
re*Matches the preceding subexpression zero or more times.
re+Matches the preceding subexpression one or more times.
re?Matches the preceding subexpression zero or one time.
re{ n}Matches the preceding subexpression exactly n times.
re{ n,}Matches the preceding subexpression n times or more.
re{ n, m}Matches the preceding subexpression at least n times and at most m times.
a| bMatches a or b.
(re)Groups a regular expression and remembers the matched text.
(?imx)Temporarily turns on the i, m, or x option within the regular expression. If inside parentheses, it only affects the part within the parentheses.
(?-imx)Temporarily turns off the i, m, or x option within the regular expression. If inside parentheses, it only affects the part within the parentheses.
(?: re)Groups a regular expression but does not remember the matched text.
(?imx: re)Temporarily turns on i, m, or x options within parentheses.
(?-imx: re)Temporarily turns off i, m, or x options within parentheses.
(?#...)Comment.
(?= re)Specifies a position using a pattern. Has no range.
(?! re)Specifies a position using a negated pattern. Has no range.
(?> re)Matches an independent pattern without backtracking.
\wMatches a word character.
\WMatches a non-word character.
\sMatches a whitespace character. Equivalent to [\t\n\r\f].
\SMatches a non-whitespace character.
\dMatches a digit. Equivalent to [0-9].
\DMatches a non-digit.
\AMatches the beginning of a string.
\ZMatches the end of the string. If a newline exists, it matches only before the newline.
\zMatches the end of the string.
\GMatches the point where the last match completed.
\bMatches a word boundary when outside brackets; matches a backspace (0x08) when inside brackets.
\BMatches a non-word boundary.
\n, \t, etc.Matches newline, carriage return, tab, etc.
\1...\9Matches the nth grouped subexpression.
\10If it has been matched, matches the nth grouped subexpression. Otherwise, it refers to the octal representation of a character encoding.

Regular Expression Examples

Characters

ExamplesDescription
/ruby/Matches "ruby"
¥Matches the Yen sign. Ruby 1.9 and Ruby 1.8 support multiple characters.

Character Classes

ExamplesDescription
/[Rr]uby/ Matches "Ruby" or "ruby"
/rub[ye]/ Matches "ruby" or "rube"
/[aeiou]/Matches any lowercase vowel
/[0-9]/ Matches any digit, same as /
/[a-z]/Matches any lowercase ASCII letter
/[A-Z]/Matches any uppercase ASCII letter
/[a-zA-Z0-9]/Matches any character in the brackets
/[^aeiou]/ Matches any character that is not a lowercase vowel
/[^0-9]/Matches any non-digit character

Special Character Classes

ExamplesDescription
/./ Matches any character except newline
/./m Matches newline as well in multiline mode
/\d/Matches a digit, equivalent to /[0-9]/
/\D/ Matches a non-digit, equivalent to /[^0-9]/
/\s/Matches a whitespace character, equivalent to /[ \t\r\n\f]/
/\S/ Matches a non-whitespace character, equivalent to /[^ \t\r\n\f]/
/\w/ Matches a word character, equivalent to /[A-Za-z0-9_]/
/\W/Matches a non-word character, equivalent to /[^A-Za-z0-9_]/

Repetition

ExamplesDescription
/ruby?/ Matches "rub" or "ruby". The y is optional.
/ruby*/ Matches "rub" plus zero or more y's.
/ruby+/Matches "rub" plus one or more y's.
/\d{3}/Matches exactly 3 digits.
/\d{3,}/Matches 3 or more digits.
/\d{3,5}/Matches 3, 4, or 5 digits.

Non-greedy Repetition

This matches the smallest number of repetitions.

ExamplesDescription
/<.*>/Greedy repetition: matches "<ruby>perl>"
/<.*?>/ Non-greedy repetition: matches "<ruby>" in "<ruby>perl>"

Grouping with Parentheses

ExamplesDescription
/\D\d+/ Without grouping: + repeats \d
/(\D\d)+/ With grouping: + repeats the \D\d pair
/([Rr]uby(, )?)+/Matches "Ruby", "Ruby, ruby, ruby", etc.

Backreferences

This matches a previously matched group again.

ExamplesDescription
/([Rr])uby&\1ails/Matches ruby&rails or Ruby&Rails
/(['"])(?:(?!\1).)*\1/Single-quoted or double-quoted strings. \1 matches the characters matched by the first group, \2 matches those matched by the second group, and so on.

Alternation

ExamplesDescription
/ruby|rube/Matches "ruby" or "rube"
/rub(y|le)/Matches "ruby" or "ruble"
/ruby(!+|\?)/ "ruby" followed by one or more ! or a ?

Anchor

This requires specifying a match position.

ExamplesDescription
/^Ruby/Matches a string or line beginning with "Ruby"
/Ruby$/ Match a string or line ending with "Ruby"
/\ARuby/ Match a string beginning with "Ruby"
/Ruby\Z/Match a string ending with "Ruby"
/\bRuby\b/Match "Ruby" at a word boundary
/\brub\B/\B is a non-word boundary: matches "rub" in "rube" and "ruby", but not a standalone "rub"
/Ruby(?=!)/Match "Ruby" if followed by an exclamation mark
/Ruby(?!!)/ Match "Ruby" if not followed by an exclamation mark

Special Syntax for Parentheses

ExamplesDescription
/R(?#comment)/ Match "R". All remaining characters are comments.
/R(?i)uby/ Case-insensitive when matching "uby".
/R(?i:uby)/ Same as above.
/rub(?:y|le))/Only group, no \1 backreference

Search and Replace

subandgsuband their substitution variablessub!andgsub!are important string methods when using regular expressions.

All these methods perform search and replace operations using regular expression patterns.subandsub!Replaces the first occurrence of the pattern,gsubandgsub!Replaces all occurrences of the pattern.

subandgsubReturns a new string, leaving the original string unmodified, whilesub!andgsub!modifies the string on which they are called.

Examples

#!/usr/bin/ruby # -*- coding: UTF-8 -*- phone = "138-3453-1111 #This is a phone number" #Remove Ruby comments phone = phone.sub!(/#.*$/, "") puts "Phone number : #{phone}" #Remove characters other than digits phone = phone.gsub!(/\D/, "") puts "Phone number : #{phone}"

Try it »

The output of the above example is:

电话号码 : 138-3453-1111 
电话号码 : 13834531111

Examples

#!/usr/bin/ruby # -*- coding: UTF-8 -*- text = "rails is rails, Ruby on Rails is a very good Ruby framework" #Replace all "rails" with "Rails" text.gsub!("rails", "Rails") #Capitalize the first letter of each occurrence of the word "Rails" text.gsub!(/\brails\b/, "Rails") puts "#{text}"

Try it »

The output of the above example is:

Rails 是 Rails,  Ruby on Rails 非常好的 Ruby 框架
Other extensions