Concepts / Escape Sequences in Regular Expressions

Escape Sequences in Regular Expressions

Square brackets [...] define a set of acceptable characters for a single position in a regex pattern.

  • Programming

From Dirty Text to Precise Matches

Regular expressions become more useful when they describe not only what must appear, but also what is allowed at the boundaries of a match. Suppose text contains source@collab.sakaiproject.org>; . A pattern that only looks for the @ symbol can capture unwanted characters attached to the email address. Character sets let you require an acceptable first and last character, so the match stops at the intended data instead of including surrounding punctuation.

What do you think happens?

What should the pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] match inside source@collab.sakaiproject.org>;?

  • source@collab.sakaiproject.org
  • source@collab.sakaiproject.org>;
  • source@
  • The entire surrounding sentence
Reveal answer

Answer: source@collab.sakaiproject.org

The first character set requires the match to begin with a letter or digit, and the final character set requires it to end with a letter. The > and ; characters cannot satisfy the final [a-zA-Z], so they remain outside the match.

One Position, Several Allowed Characters

Square brackets define a character set. A character set describes the acceptable characters for one position in a regular expression. The set [a-z] allows any single character from a through z. The set [a-zA-Z0-9] allows one lowercase letter, one uppercase letter, or one digit. It does not describe a multi-character string as a single choice; it describes one position whose value may be any one character from the listed range.

allowsallowsallowsOne position[a-zA-Z0-9]a-zlowercase letterA-Zuppercase letter0-9digit
Which individual characters can occupy one position matched by [a-zA-Z0-9]?

[a-zA-Z0-9]

You have already seen an implicit character set in the escape sequence \S. It matches any non-whitespace character. Square brackets provide an explicit way to list the characters or ranges that are acceptable.

How Quantifiers Extend a Match

The quantifiers * and + control repetition, but each applies only to the character immediately to its left. The quantifier * means zero or more repetitions. The quantifier + means one or more repetitions. In \S*, the asterisk applies to \S, not to the entire regular expression and not to the @ symbol. In the same way, \S+ would require one or more non-whitespace characters.

permitspermitsrequires\S*zero or more0non-whitespace characters1 or morenon-whitespace characters\S+one or more1 or morenon-whitespace characters
What does each quantifier allow for the character immediately before it?
ExpressionMeaningMinimum repetitions
\S*Zero or more non-whitespace characters0
\S+One or more non-whitespace characters1

Why the Email Pattern Uses an Asterisk

Choose between [a-zA-Z0-9]\S*@ and [a-zA-Z0-9]\S+@ when the pattern should allow the address a@b.c.

Count the starting character: [a-zA-Z0-9] already matches the first character, such as a.

Check the gap before @: For a@b.c, there are no additional characters between the initial a and the @ symbol.

Apply the quantifier rule: \S* allows zero additional non-whitespace characters, while \S+ requires at least one additional character.

[a-zA-Z0-9]\S*@ allows a@b.c; using \S+ in that position would require at least two non-whitespace characters before the @ when combined with the initial character set.

Building the Email Pattern

[a-zA-Z0-9]\S*@\S*[a-zA-Z]

The pattern has five parts. First, [a-zA-Z0-9] matches exactly one letter or digit and prevents the match from starting with a symbol such as <. Next, \S* matches zero or more non-whitespace characters. The @ symbol matches itself literally. A second \S* matches zero or more non-whitespace characters after the @. Finally, [a-zA-Z] matches exactly one letter, ensuring that the match ends with a letter rather than a symbol such as > or ;.

thenthenthenthen[a-zA-Z0-9]one letter or digit\S*zero or more non-whitespace@literal symbol\S*zero or more non-whitespace[a-zA-Z]one letter
How do character sets, the literal @ symbol, and quantifiers connect to match the intended parts of an email address?

Filtering the Pattern Boundaries

Consider source@collab.sakaiproject.org>;. The pattern can consume non-whitespace characters through the address, but the final [a-zA-Z] limits the endpoint. When the engine reaches >, that character cannot satisfy the final character set. The match therefore ends at the last acceptable letter, g in org, leaving >; outside the matched text.

boundary filteringleaves outside<source@collab.sakaiproject.org>;surrounding text andpunctuation>;outside the matchsource@collab.sakaiproject.orgmatched email
How do the starting and ending character sets remove unwanted boundary characters while preserving the email address?

This is a boundary-finding effect. The character sets specify exactly what is acceptable at the two ends, while the regular expression engine finds the stopping point that satisfies those requirements. The first set rejects a leading symbol such as <, and the final set rejects trailing symbols such as > or ;.

Reading findall Results

inputreturnsInput texttext containing matchesre.findall()find matchesPython listmatched strings
What does the input text become after re.findall() identifies matching email strings?

The result is not one combined string. re.findall() returns a Python list, and each found match is a string element in that list. If the pattern matches more than once in a single line, the list contains multiple strings.

Common Pattern Mistakes

  • Treating [a-zA-Z0-9] as a multi-character string

    Square brackets define acceptable characters for one position. The ranges represent possible individual characters.

    Fix: Read the set as one position that may contain one lowercase letter, uppercase letter, or digit.

  • Applying * to more than the immediately preceding character

    The quantifier applies only to the character immediately to its left, which is \S.

    Fix: Parse \S* as zero or more non-whitespace characters followed by a literal @.

  • Using + after the initial character set without checking the minimum length

    The initial character set already matches one non-whitespace character. Adding + requires at least one more non-whitespace character before @.

    Fix: Use \S* when no additional character is required between the initial character and @.

  • Leaving the boundaries unrestricted

    The pattern does not require an acceptable letter or digit at the start or a letter at the end, so unwanted boundary characters may be included.

    Fix: Add [a-zA-Z0-9] at the start and [a-zA-Z] at the end when those boundaries are required.

  • Expecting findall() to return one combined value

    re.findall() returns a Python list, with each match represented as a string element.

    Fix: Interpret the result as a list of matched strings.

Practice: Assemble the Match

MEDIUM

Explain each part of [a-zA-Z0-9]\S*@\S*[a-zA-Z], then determine which portion of source@collab.sakaiproject.org>; is returned as the match.

Hints
  • Identify what the first character set permits.
  • Remember that each * applies only to the immediately preceding \S.
  • Check the final character against [a-zA-Z].
MEDIUM

Compare \S* and \S+ in the portion before the @ symbol. Which one permits the starting character set to be the only character before @, and why?

Hints
  • The first character set already consumes one character.
  • Focus on the minimum number of additional non-whitespace characters.

Key Takeaways

  1. Square brackets define a set of acceptable characters for one position in a regular expression.
  2. Ranges such as [a-z] and [a-zA-Z0-9] represent groups of possible individual characters.
  3. The quantifiers * and + apply only to the character immediately to their left; * permits zero or more repetitions, while + requires one or more.
  4. In [a-zA-Z0-9]\S*@\S*[a-zA-Z], the boundary character sets filter unwanted symbols while \S* provides flexibility in the middle.
  5. re.findall() returns a Python list whose elements are the matched strings.

Key Takeaways

  • Character sets use square brackets to define acceptable characters for a single position.
  • The * and + quantifiers repeat only the immediately preceding character or escape sequence.
  • The email pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] uses flexible middle sections and restricted boundaries.
  • Boundary character sets prevent punctuation such as <, >, and ; from becoming part of the extracted match.
  • re.findall() returns matched strings as elements of a Python list.