Building Complex Patterns with Alternation and Grouping
Raw strings (r'...') are mandatory for regex patterns because they prevent Python from interpreting backslashes as escape sequences, allowing the regex engine to receive the correct pattern.
A Backslash Has Two Interpreters
A regular-expression pattern containing a backslash passes through two stages. First, Python interprets the string literal. Then the regular-expression engine receives the resulting pattern. Python uses backslashes for escape sequences such as \n and \t, while a regular expression uses backslashes for patterns such as \S. A raw string, written with the r prefix, prevents Python from interpreting the backslashes before the regular-expression engine receives the pattern.
How \S+ Consumes Text
The \S character class matches one non-whitespace character. Adding + changes the pattern to \S+, which matches one or more consecutive non-whitespace characters. The match continues across adjacent non-whitespace characters and stops when it encounters whitespace or reaches the end of the string.
Predict the \S+ matches
Determine the substrings matched by \S+ in the text alpha beta\tgamma.
Start at alpha: The characters in alpha are consecutive and non-whitespace, so they form the first match.
Reach the space: The space is whitespace, so the first match stops before it.
Continue at beta: beta forms the next consecutive non-whitespace run.
Reach the tab: The tab is whitespace, so the second match stops before it.
Finish at gamma: gamma contains consecutive non-whitespace characters and continues to the end of the input.
The matches are alpha, beta, and gamma.
['alpha', 'beta', 'gamma']Finding Email-Like Substrings
The pattern \S+@\S+ combines three requirements: one or more non-whitespace characters before the @ symbol, the @ symbol itself, and one or more non-whitespace characters after it. It extracts email-like substrings from unstructured text. This pattern does not establish that a substring is a complete or valid email address; it only applies the stated non-whitespace requirements around @.
['sam@example.com', 'lee@school.org.']How findall Scans the Input
re.findall() scans the input from left to right. At each position, it tests whether the pattern matches there. When it finds a match, it adds the matched substring to a list and continues from the position after that match. The scan repeats until the end of the input, producing a list containing all matches.
import re line = "Send to sam@example.com and lee@school.org" matches = re.findall(r"\S+@\S+", line) if len(matches) > 0: print(matches)
Diagnosing Missing Matches
Writing a backslash-containing pattern as an ordinary string without considering Python's string interpretation.
Python processes backslashes in regular strings before the regular-expression engine receives the pattern, which can create a mismatch between the intended pattern and the pattern used.
Fix:
Use the raw-string form r"\S+@\S+".Expecting \S+ to cross whitespace.
\S+ stops when it encounters whitespace.
Fix:
Predict separate matches on either side of the whitespace.Assuming \S+@\S+ matches every string that looks email-like to a person.
The pattern requires at least one non-whitespace character on both sides of @.
Fix:
Check both sides of @ against the pattern's explicit requirements.Accessing an item from an empty result list.
re.findall() returns an empty list when no match is found, so indexing it can cause IndexError.
Fix:
Check that the list is not empty, for example with len(matches) > 0, before accessing an item.
The pattern \S+@\S+ requires at least one non-whitespace character before @ and at least one after it. Therefore, an input with a missing side of @ does not satisfy this pattern. When no substring satisfies the pattern, re.findall() returns an empty list, which must be handled before individual list elements are accessed.
Practice the Scan
What do you think happens?
What list will re.findall(r"\S+", "one two\tthree") return?
Reveal answer
Answer: ['one', 'two', 'three']
\S+ consumes consecutive non-whitespace characters. The space and tab divide the input into three separate runs.
Write a short Python expression that extracts all email-like substrings from the text "Names: pat@example.net, robin@work.org". Then explain why a raw string is used for the pattern and what happens if the text contains no match.
Hints
- Use re.findall().
- Use the pattern r"\S+@\S+".
- Remember that the result is a list and can be empty.
Reliable Pattern Habits
- Write backslash-containing regular-expression patterns as raw strings.
- Read \S+ as one or more consecutive non-whitespace characters.
- Mark spaces, tabs, and line breaks as possible boundaries between \S+ matches.
- Use \S+@\S+ when the task requires non-whitespace text on both sides of @.
- Treat the result of re.findall() as a list that may contain zero, one, or multiple matches.
- Check that the list is not empty before accessing an individual element.
Key Takeaways
- Raw strings preserve backslashes for the regular-expression engine instead of allowing Python to interpret them first.
- \S matches a non-whitespace character, while \S+ matches a consecutive run of one or more non-whitespace characters.
- Whitespace divides \S+ matches, and a match also stops at the end of the input.
- The pattern \S+@\S+ extracts email-like substrings with non-whitespace text on both sides of @.
- re.findall() returns a list of all matches, so check that the list is not empty before indexing it.