Regular Expression Fundamentals
Parentheses in regex patterns create capturing groups that extract specific substrings from a match.
From HTML to URLs
Suppose an HTML document contains several links and your goal is to collect only their URL values. A regular expression can describe the surrounding HTML, while a capturing group identifies the part you actually want to extract. The important mechanism is the relationship between the complete match and the parenthesized substring inside it.
What do you think happens?
A pattern contains a larger match with a parenthesized URL portion. What should re.findall() return when used with that pattern?
Reveal answer
Answer: Only the text captured by the parentheses
In this extraction pattern, re.findall() returns the captured groups rather than the entire pattern match.
Tracing a URL Pattern
A regular expression can match more text than you want to return. Parentheses mark a capturing group, so the text matched inside those parentheses is isolated from the rest of the pattern. In a URL pattern, the surrounding text can identify an href attribute while the group identifies the URL value inside the quotation marks.
The characters .*? are especially important in this pattern. The dot and asterisk describe a span of text, while the question mark makes that span non-greedy. The non-greedy form stops at the first closing quote instead of continuing across later text and matching too much.
Capturing the Substring
Separating the Match from the Result
Use a pattern with one capturing group to extract the URL from an HTML link.
Describe the surrounding text: The pattern begins by matching the href attribute marker and ends by matching its closing quote.
Create the group: The parentheses around .*? identify the substring between those quotes as the captured value.
Control the span: The non-greedy quantifier makes the group stop at the first closing quote rather than consuming too much text.
Read the extraction: When re.findall() processes the pattern, it returns the captured URL values instead of the complete href fragments.
The surrounding match locates the attribute, and the capturing group supplies the URL value returned by re.findall().
import re html = '<a href="https://example.test">Example</a>' pattern = r'href="(.*?)"' urls = re.findall(pattern, html) print(urls)
The findall Data Flow
The extraction process has three useful stages. First, HTML content is supplied as the text to search. Next, the regular expression locates the relevant attribute and captures the substring inside it. Finally, re.findall() collects the captured values into a list. Its behavior is especially useful here because it returns the captured groups, not the entire pattern match.
["https://one.test", "https://two.test"]Debugging Against Actual HTML
A failed extraction should be debugged by comparing the pattern with the actual HTML, not by staring at the pattern in isolation. Trace the match from left to right. Check the attribute marker, check the quote character, check the captured span, and check the closing quote. A mismatch in the surrounding HTML can prevent the capturing group from being reached at all.
[]When the result is empty, inspect the exact HTML text first. Confirm that the pattern's surrounding text appears in the same form, then verify that the quoted portion is where the pattern expects it. If the pattern matches too much, inspect whether .*? was used; the non-greedy form is essential for stopping at the first closing quote.
SSL Contexts for Requests
URL extraction may begin with fetching HTML from an HTTPS site. If the site has an invalid certificate, certificate verification must be disabled in the fetching code by using an SSL context. This setting concerns the HTTPS request that obtains the HTML; it does not change what the regular expression captures after the HTML has been received.
- Use an SSL context when the request must handle a site with an invalid certificate.
- Obtain the HTML content through that configured request.
- Inspect the actual HTML returned by the request.
- Apply the regular expression and verify the captured URL values.
Common Extraction Mistakes
Using a greedy span when extracting a quoted URL
The match can continue farther than the first closing quote and consume too much text.
Fix:
Use the non-greedy quantifier .*? so the URL extraction stops at the first closing quote.Expecting re.findall() to return the complete HTML match
With a capturing group, re.findall() returns the captured group values rather than the entire pattern match.
Fix:
Read the returned list as the extracted substrings isolated by the parentheses.Debugging only the regular expression
The pattern can fail before the capturing group is reached when the surrounding HTML does not match.
Fix:
Compare each part of the pattern with the actual HTML returned by the request.Treating an HTTPS certificate problem as a regex problem
There is no correct HTML text for the regex to inspect until the request is handled.
Fix:
Configure an SSL context for the invalid-certificate case, then debug extraction against the HTML that was actually received.
Practice the Trace
Given the HTML text <a href="https://one.test">One</a> and the pattern href="(.*?)", identify the complete portion matched by the pattern, the substring captured by the parentheses, and the value that re.findall() places in its result.
Hints
- Separate the surrounding href text from the text inside the parentheses.
- The non-greedy group stops at the first closing quote.
- re.findall() returns the captured URL value rather than the complete href fragment.
A URL extraction returns an empty list. The request used an SSL context for a site with an invalid certificate, but the returned HTML uses single quotes around href values while the pattern expects double quotes. Trace the failure and state which part must be investigated first.
Hints
- First separate the request stage from the matching stage.
- Confirm that HTML was received before examining the pattern.
- Compare the quote characters in the actual HTML with the quote characters expected by the pattern.
Key Takeaways
- Parentheses create capturing groups that isolate specific substrings from a larger regular expression match.
- The non-greedy quantifier .*? is essential when a URL must stop at the first closing quote.
- re.findall() returns captured group values, making it useful for collecting URLs from HTML.
- Reliable debugging means tracing the pattern against the actual HTML rather than inspecting the pattern alone.
- An SSL context can be used to disable certificate verification when fetching from a site with an invalid certificate.
Key Takeaways
- Capturing groups use parentheses to isolate the substring that matters.
- Use .*? for quoted URL extraction so matching stops at the first closing quote.
- re.findall() returns captured values instead of the entire pattern match.
- Debug extraction by comparing every pattern component with the actual HTML.
- Use an SSL context to handle certificate verification when fetching from sites with invalid certificates.