Anchors and Boundaries in Regular Expressions
Square brackets [...] define a set of acceptable characters for a single position in a regex pattern.
From Dirty Text to Clean Matches
Suppose text contains the string source@collab.sakaiproject.org>;. The email address is present, but extra characters cling to its boundary. A pattern that only looks for the @ symbol can capture those unwanted characters. The central idea in this article is to make the beginning and ending of a match more specific: require an alphanumeric character at the start and a letter at the end.
One Position, Several Acceptable Characters
Square brackets define a character set. A character set describes the alternatives allowed at one position in a regular expression, so it matches exactly one character from the set. For example, [A-Z] accepts one uppercase letter. A range such as [a-z] represents every character from a through z. Combining ranges gives [a-zA-Z0-9], which accepts one lowercase letter, uppercase letter, or digit.
A character set does not automatically match a whole word. Without a quantifier, [a-zA-Z0-9] matches one character only.
What the Quantifier Repeats
The quantifiers * and + apply only to the character immediately to their left. The quantifier * means zero or more occurrences, while + means one or more occurrences. In \S*, the asterisk applies to \S, so the pattern can consume zero or more non-whitespace characters. In a pattern containing a character set followed by a quantifier, the quantifier applies to that character set as the immediately preceding unit.
Why the Email Pattern Uses an Asterisk
Choose between [a-zA-Z0-9]\S*@ and [a-zA-Z0-9]\S+@ when the pattern must allow the email address a@b.c.
Count the initial character: [a-zA-Z0-9] already matches the initial a.
Check the gap before @: There are no additional characters between the initial a and the @ symbol, so the middle part must be allowed to match zero characters.
Choose the quantifier: \S* allows zero or more non-whitespace characters. \S+ would require at least one additional non-whitespace character after the initial character.
[a-zA-Z0-9]\S*@ allows the stated short email form, while [a-zA-Z0-9]\S+@ would exclude it.
Filtering the Email Boundaries
The source pattern for extracting an email address is [a-zA-Z0-9]\S*@\S*[a-zA-Z]. It has five functional parts. The first character set requires the match to begin with a letter or digit. The first \S* allows zero or more non-whitespace characters before the @ symbol. The @ matches itself literally. The second \S* allows zero or more non-whitespace characters after the @ symbol. The final [a-zA-Z] requires the match to end with a letter.
When the engine reaches the > character after org, that character cannot satisfy the final [a-zA-Z]. The engine therefore backtracks to the last letter it found, the g in org, and the match stops there. The ending character set acts as a precise stopping condition rather than allowing every non-whitespace character to remain in the result.
Reading the Extraction Result
The re.findall() function returns a Python list. Each match appears as a string element in that list. If a pattern finds multiple matches in one line, the list contains multiple strings in the order returned by the search.
For the boundary-filtering pattern, a match from source@collab.sakaiproject.org>; is represented as the string source@collab.sakaiproject.org rather than including the trailing > or ;. Multiple matches are represented as separate string elements in the returned list.
Mistakes in Character-Set Patterns
Treating a character set as if it matches a whole word
A character set defines acceptable characters for one position and matches exactly one character unless a quantifier changes that behavior.
Fix:
Use a quantifier when the pattern needs to consume additional characters, and remember that the quantifier applies only to the item immediately before it.Using + after the initial character set in the email pattern
The initial character set already accounts for one non-whitespace character. The + then requires at least one more non-whitespace character before @, excluding a@b.c.
Fix:
Use \S* when zero additional characters must be allowed after the initial alphanumeric character.Allowing \S at both boundaries without restrictions
\S matches non-whitespace characters, including boundary symbols such as <, >, or ;.
Fix:
Use [a-zA-Z0-9] at the start and [a-zA-Z] at the end when those are the required boundary characters.Assuming the @ symbol alone produces clean email data
Extra punctuation can cling to the email address in the source text.
Fix:
Make the acceptable starting and ending characters explicit.
Build and Check a Precise Pattern
Construct the five-part pattern for an email address that starts with a letter or digit, allows zero or more non-whitespace characters before and after @, and ends with a letter. Then explain why the pattern stops before a trailing > or ;.
Hints
- Begin with [a-zA-Z0-9].
- Place \S* before the literal @ and another \S* after it.
- Finish with [a-zA-Z].
- Check the final character of source@collab.sakaiproject.org>; against the ending character set.
What do you think happens?
What value should the pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] extract from source@collab.sakaiproject.org>;
Reveal answer
Answer: source@collab.sakaiproject.org
The final [a-zA-Z] requires the match to end with a letter. The > and ; characters cannot satisfy it, so the match stops at the g in org.
Key Takeaways
- Square brackets define a set of acceptable characters for one position in a regular expression.
- Ranges such as [a-z] represent all characters in the stated range, and [a-zA-Z0-9] accepts one letter or digit.
- The quantifiers * and + apply only to the character or pattern unit immediately to their left; * allows zero or more, while + allows one or more.
- In [a-zA-Z0-9]\S*@\S*[a-zA-Z], the boundary character sets restrict the beginning and ending characters while \S* provides flexibility in the middle.
- re.findall() returns matches as strings inside a Python list.
Key Takeaways
- Character sets describe the acceptable character at a single position.
- Quantifiers change how many times the immediately preceding character or set can match.
- Boundary character sets prevent unwanted punctuation from entering extracted email addresses.
- The pattern [a-zA-Z0-9]\S*@\S*[a-zA-Z] uses flexible middle sections and restrictive edges.
- re.findall() packages each match as a string in a Python list.