Character Classes and Quantifiers in Regex
Raw strings (r'...') are mandatory for regex patterns because they prevent Python from interpreting backslashes as escape sequences, allowing the regex engine to receive the correct pattern.
From Python Source to Regex Engine
A regular expression pattern may contain backslashes, but Python processes the pattern before the regular expression engine receives it. In an ordinary Python string, Python treats backslashes as special characters used in escape sequences such as \n and \t. A raw string, written with the r prefix, prevents Python from interpreting those backslashes first. This allows the regex engine to receive the pattern you intended, such as \S+.
Reading \S and + Together
The character class \S matches one non-whitespace character. The plus quantifier means one or more, so \S+ matches one or more consecutive non-whitespace characters. The match continues across adjacent non-whitespace characters and stops when it reaches whitespace or the end of the string.
What do you think happens?
What substrings will r'\S+' match in the text Send mail today?
Reveal answer
Answer: Send, mail, and today as three matches
Each word is a consecutive run of non-whitespace characters. The spaces stop one \S+ match and separate it from the next.
['Send', 'mail', 'today']Scanning with findall
re.findall() scans the input from left to right. At each position, it tests whether the pattern matches. When it finds a match, it adds the matched substring to a list and continues scanning from the position after that match. The result is the complete list of matches found by the scan.
Extracting Non-Whitespace Tokens
Use r'\S+' with re.findall() on the text 'Status: ready now'.
Construct the pattern: The raw string r'\S+' gives the regex engine a pattern that means one or more consecutive non-whitespace characters.
Scan from the left: The scan first finds Status:, then stops at the space.
Continue after whitespace: The scan finds ready and then now as separate non-whitespace runs.
The returned list is ['Status:', 'ready', 'now'].
Finding Email-Like Text
The pattern \S+@\S+ requires at least one non-whitespace character before the @ symbol and at least one non-whitespace character after it. Combined with re.findall(), it extracts email-like substrings from surrounding prose. Text that does not satisfy those requirements remains unmatched.
import re text = 'Contact alice@example.com or bob@example.org.' addresses = re.findall(r'\S+@\S+', text) print(addresses)
Debugging Empty Results
Leaving out the raw-string prefix.
Python processes backslashes in regular strings before the regex engine receives the pattern, which can create a mismatch between the intended pattern and the pattern used.
Fix:
Write pattern = r'\S+'.Expecting \S+ to cross whitespace.
\S matches non-whitespace, and the + quantifier stops when whitespace is reached.
Fix:
Expect separate matches for 'red' and 'blue'.Assuming every email-like string satisfies \S+@\S+.
The pattern requires at least one non-whitespace character immediately before @, but the space breaks that requirement.
Fix:
Compare the input with the pattern and check that non-whitespace text appears on both sides of @.Accessing the first result without checking whether a match exists.
re.findall() returns an empty list when there are no matches, so accessing an element that is not present can cause IndexError.
Fix:
Check that the returned list is not empty before accessing individual elements.
When a match is missing, inspect the problem in stages: confirm that the pattern is a raw string, check what \S and + require, compare those requirements with the input text, and verify that the result list contains an item before reading one.
Practice the Prediction
Predict the result of re.findall(r'\S+', 'A\tshort\nline') before checking it. Identify each non-whitespace run and explain what causes one match to stop and the next one to begin.
Hints
- Treat the tab and line break as whitespace separators.
- The + quantifier consumes consecutive non-whitespace characters only.
- List the substrings in their left-to-right order.
Use re.findall(r'\S+@\S+', text) on a line that contains one valid email-like substring and one string with a space before @. Predict which substring appears in the returned list, then add a length check before processing the result.
Hints
- The pattern needs non-whitespace text immediately before and after @.
- re.findall() returns all matches in a list.
- Check len(matches) > 0 before accessing an individual result.
Key Takeaways
- Use raw strings such as r'\S+' so Python does not interpret regex backslashes before the regex engine receives them.
- \S matches one non-whitespace character, while \S+ matches a consecutive run of one or more non-whitespace characters.
- re.findall() scans from left to right and returns all matches in a list.
- The pattern \S+@\S+ extracts email-like substrings with non-whitespace text on both sides of @.
- Check that the result list is not empty before accessing an individual match.
Key Takeaways
- Raw strings protect backslashes in regex patterns from Python's string interpretation.
- \S+ consumes each consecutive non-whitespace run and stops at whitespace or the end of the input.
- re.findall() collects every match while scanning from left to right.
- \S+@\S+ is useful for extracting email-like substrings, but it only enforces its stated non-whitespace requirements.
- Empty results must be checked before individual list elements are accessed.