Escape Sequences in Regular Expressions
Square brackets [...] define a set of acceptable characters for a single position in a regex pattern.
From Dirty Text to Precise Matches
Regular expressions become more useful when they describe not only what must appear, but also what is allowed at the boundaries of a match. Suppose text contains source@collab.sakaiproject.org>; . A pattern that only looks for the @ symbol can capture unwanted characters attached to the email address. Character sets let you require an acceptable first and last character, so the match stops at the intended data instead of including surrounding punctuation.
What do you think happens?
What should the pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] match inside source@collab.sakaiproject.org>;?
Reveal answer
Answer: source@collab.sakaiproject.org
The first character set requires the match to begin with a letter or digit, and the final character set requires it to end with a letter. The > and ; characters cannot satisfy the final [a-zA-Z], so they remain outside the match.
One Position, Several Allowed Characters
Square brackets define a character set. A character set describes the acceptable characters for one position in a regular expression. The set [a-z] allows any single character from a through z. The set [a-zA-Z0-9] allows one lowercase letter, one uppercase letter, or one digit. It does not describe a multi-character string as a single choice; it describes one position whose value may be any one character from the listed range.
[a-zA-Z0-9]
You have already seen an implicit character set in the escape sequence \S. It matches any non-whitespace character. Square brackets provide an explicit way to list the characters or ranges that are acceptable.
How Quantifiers Extend a Match
The quantifiers * and + control repetition, but each applies only to the character immediately to its left. The quantifier * means zero or more repetitions. The quantifier + means one or more repetitions. In \S*, the asterisk applies to \S, not to the entire regular expression and not to the @ symbol. In the same way, \S+ would require one or more non-whitespace characters.
| Expression | Meaning | Minimum repetitions |
|---|---|---|
| \S* | Zero or more non-whitespace characters | 0 |
| \S+ | One or more non-whitespace characters | 1 |
Why the Email Pattern Uses an Asterisk
Choose between [a-zA-Z0-9]\S*@ and [a-zA-Z0-9]\S+@ when the pattern should allow the address a@b.c.
Count the starting character: [a-zA-Z0-9] already matches the first character, such as a.
Check the gap before @: For a@b.c, there are no additional characters between the initial a and the @ symbol.
Apply the quantifier rule: \S* allows zero additional non-whitespace characters, while \S+ requires at least one additional character.
[a-zA-Z0-9]\S*@ allows a@b.c; using \S+ in that position would require at least two non-whitespace characters before the @ when combined with the initial character set.
Building the Email Pattern
[a-zA-Z0-9]\S*@\S*[a-zA-Z]
The pattern has five parts. First, [a-zA-Z0-9] matches exactly one letter or digit and prevents the match from starting with a symbol such as <. Next, \S* matches zero or more non-whitespace characters. The @ symbol matches itself literally. A second \S* matches zero or more non-whitespace characters after the @. Finally, [a-zA-Z] matches exactly one letter, ensuring that the match ends with a letter rather than a symbol such as > or ;.
Filtering the Pattern Boundaries
Consider source@collab.sakaiproject.org>;. The pattern can consume non-whitespace characters through the address, but the final [a-zA-Z] limits the endpoint. When the engine reaches >, that character cannot satisfy the final character set. The match therefore ends at the last acceptable letter, g in org, leaving >; outside the matched text.
This is a boundary-finding effect. The character sets specify exactly what is acceptable at the two ends, while the regular expression engine finds the stopping point that satisfies those requirements. The first set rejects a leading symbol such as <, and the final set rejects trailing symbols such as > or ;.
Reading findall Results
The result is not one combined string. re.findall() returns a Python list, and each found match is a string element in that list. If the pattern matches more than once in a single line, the list contains multiple strings.
Common Pattern Mistakes
Treating [a-zA-Z0-9] as a multi-character string
Square brackets define acceptable characters for one position. The ranges represent possible individual characters.
Fix:
Read the set as one position that may contain one lowercase letter, uppercase letter, or digit.Applying * to more than the immediately preceding character
The quantifier applies only to the character immediately to its left, which is \S.
Fix:
Parse \S* as zero or more non-whitespace characters followed by a literal @.Using + after the initial character set without checking the minimum length
The initial character set already matches one non-whitespace character. Adding + requires at least one more non-whitespace character before @.
Fix:
Use \S* when no additional character is required between the initial character and @.Leaving the boundaries unrestricted
The pattern does not require an acceptable letter or digit at the start or a letter at the end, so unwanted boundary characters may be included.
Fix:
Add [a-zA-Z0-9] at the start and [a-zA-Z] at the end when those boundaries are required.Expecting findall() to return one combined value
re.findall() returns a Python list, with each match represented as a string element.
Fix:
Interpret the result as a list of matched strings.
Practice: Assemble the Match
Explain each part of [a-zA-Z0-9]\S*@\S*[a-zA-Z], then determine which portion of source@collab.sakaiproject.org>; is returned as the match.
Hints
- Identify what the first character set permits.
- Remember that each * applies only to the immediately preceding \S.
- Check the final character against [a-zA-Z].
Compare \S* and \S+ in the portion before the @ symbol. Which one permits the starting character set to be the only character before @, and why?
Hints
- The first character set already consumes one character.
- Focus on the minimum number of additional non-whitespace characters.
Key Takeaways
- Square brackets define a set of acceptable characters for one position in a regular expression.
- Ranges such as [a-z] and [a-zA-Z0-9] represent groups of possible individual characters.
- The quantifiers * and + apply only to the character immediately to their left; * permits zero or more repetitions, while + requires one or more.
- In [a-zA-Z0-9]\S*@\S*[a-zA-Z], the boundary character sets filter unwanted symbols while \S* provides flexibility in the middle.
- re.findall() returns a Python list whose elements are the matched strings.
Key Takeaways
- Character sets use square brackets to define acceptable characters for a single position.
- The * and + quantifiers repeat only the immediately preceding character or escape sequence.
- The email pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] uses flexible middle sections and restricted boundaries.
- Boundary character sets prevent punctuation such as <, >, and ; from becoming part of the extracted match.
- re.findall() returns matched strings as elements of a Python list.