Scala Regular Expressions

Scala supports regular expressions through the scala.util.matching package'sRegexRegex class. The following example demonstrates using regular expressions to find words.Scala :

Example

import scala.util.matching.Regex

object Test {
   def main(args: Array[String]) {
      val pattern = "Scala".r
      val str = "Scala is Scalable and cool"
     
      println(pattern findFirstIn str)
   }
}

Execute the above code and the output result is:

$ scalac Test.scala 
$ scala Test
Some(Scala)

In the example, the r() method of the String class is used to construct a Regex object.

Then use the findFirstIn method to find the first match.

If you need to view all matches, you can use the findAllIn method.

You can use the mkString() method to concatenate the strings of regex match results, and can use the pipe (|) to set different patterns:

Example

import scala.util.matching.Regex

object Test {
   def main(args: Array[String]) {
      val pattern = new Regex("(S|s)cala")  // The first letter can be uppercase S or lowercase s
      val str = "Scala is scalable and cool"
     
      println((pattern findAllIn str).mkString(","))   // Use comma , to concatenate returned results
   }
}

Execute the above code and the output result is:

$ scalac Test.scala 
$ scala Test
Scala,scala

If you need to replace the matched text with a specified keyword, you can usereplaceFirstIn( )method to replace the first match, usereplaceAllIn( )method to replace all matches, the example is as follows:

Example

object Test {
   def main(args: Array[String]) {
      val pattern = "(S|s)cala".r
      val str = "Scala is scalable and cool"
     
      println(pattern replaceFirstIn(str, "Java"))
   }
}

Execute the above code and the output result is:

$ scalac Test.scala 
$ scala Test
Java is scalable and cool

Regular Expressions

Scala's regular expressions inherit Java's syntax rules, while Java largely uses Perl language rules.

The following table gives some commonly used regular expression rules:

Expression Matching rule
^ Matches the position at the start of the input string.
$ Matches the position at the end of the input string.
. Matches any single character except "\r\n".
[...] Character set. Matches any one character included. For example, "[abc]" matches "a" in "plain".
[^...] Negated character set. Matches any character not included. For example, "[^abc]" matches "p", "l", "i", "n" in "plain".
\\A Matches the start position of the input string (without multiline support)
\\z End of string (like $, but not affected by multiline processing options)
\\Z End of string or line (not affected by multiline processing options)
re* Repeat zero or more times
re+ Repeat one or more times
re? Repeat zero or one time
re{ n} Repeat exactly n times
re{ n,}
re{ n, m} Repeat between n and m times
a|b Match a or b
(re) Match re and capture the text into an automatically named group
(?: re) Match re, do not capture the matched text, and do not assign a group number to this group
(?> re) Greedy subexpression
\\w Matches letters, digits, or underscores
\\W Matches any character that is not a letter, digit, underscore, or Chinese character
\\s Matches any whitespace character, equivalent to [\t\n\r\f]
\\S Matches any character that is not a whitespace character
\\d Matches a digit, like [0-9]
\\D Matches any non-digit character
\\G Start of current search
\\n Newline character
\\b Usually a word boundary position, but if used in a character class, represents backspace
\\B Matches a position that is not the start or end of a word
\\t Tab character
\\Q Beginning quote:\Q(a+b)*3\ECan match the text "(a+b)*3".
\\E Closing quote:\Q(a+b)*3\ECan match the text "(a+b)*3".

Regular Expression Examples

Example Description
. Matches any single character except "\r\n".
[Rr]uby Matches "Ruby" or "ruby"
rub[ye] Matches "ruby" or "rube"
[aeiou] Matches lowercase letters: aeiou
[0-9] Matches any digit, like
[a-z] Matches any ASCII lowercase letter
[A-Z] Matches any ASCII uppercase letter
[a-zA-Z0-9] Matches digits, uppercase and lowercase letters
[^aeiou] Matches characters other than aeiou
[^0-9] Matches characters other than digits
\\d Matches a digit, like: [0-9]
\\D Matches a non-digit, like: [^0-9]
\\s Matches whitespace, like: [ \t\r\n\f]
\\S Matches non-whitespace, like: [^ \t\r\n\f]
\\w Matches letters, digits, underscores, like: [A-Za-z0-9_]
\\W Matches non-letters, digits, underscores, like: [^A-Za-z0-9_]
ruby? Matches "rub" or "ruby": y is optional
ruby* Matches "rub" followed by zero or more y's.
ruby+ Matches "rub" followed by one or more y's.
\\d{3} Matches exactly 3 digits.
\\d{3,} Matches 3 or more digits.
\\d{3,5} Matches 3, 4, or 5 digits.
\\D\\d+ No group: + repeats \d
(\\D\\d)+/ Group: + repeats the \D\d pair
([Rr]uby(, )?)+ Matches "Ruby", "Ruby, ruby, ruby", etc.

Note that each character in the above table uses two backslashes. This is because in Java and Scala, the backslash in strings is an escape character. So if you want to output\you need to write it in the string as\\to get a backslash. See the following example:

Example

import scala.util.matching.Regex

object Test {
   def main(args: Array[String]) {
      val pattern = new Regex("abl[ae]\\d+")
      val str = "ablaw is able1 and cool"
     
      println((pattern findAllIn str).mkString(","))
   }
}

Execute the above code and the output result is:

$ scalac Test.scala 
$ scala Test
able1
Other extensions