Working with URLs and urllib
Regular expressions provide a concise way to search for and extract patterns in HTML text, such as URLs in href attributes.
From HTML Text to URLs
HTML contains structured data embedded in text. A link commonly stores its destination in an href attribute, such as href="http://example.com". If you need every link in a page, you could search through characters manually and build substrings, but a regular expression lets you describe the pattern and ask the regex engine to find matching text.
The extraction task has two layers. First, the pattern must recognize the href attribute and its URL. Second, the pattern must identify which part of that match should be returned. The pattern href="(http[s]?://.+?)" handles both layers: the literal href=" anchors the search, while parentheses surround the URL that should be extracted.
Reading the Pattern
In href="(http[s]?://.+?)", href=" is literal text that anchors the match to the beginning of an href attribute. The expression [s]? allows either http or https. The dot and plus, .+, match the domain and path, and the question mark after the plus changes that quantifier from greedy to non-greedy.
| Pattern part | Purpose |
|---|---|
| href=" | Anchors the search to the beginning of an href attribute |
| http | Requires the URL scheme to begin with http |
| [s]? | Allows the optional s in https |
| .+? | Matches the domain and path non-greedily |
| ( ... ) | Defines the part returned as a capture group |
| " | Marks the closing quote of the attribute |
Parts of the URL-extraction pattern
The pattern must stop at the closing quote belonging to the current href attribute. That stopping behavior matters when several links occur in the same HTML text. A greedy quantifier tries to match the largest possible string, while a non-greedy quantifier tries to match the smallest possible string that still satisfies the pattern.
Tracing Two Links
Extracting Two URL Values
Suppose HTML contains two href attributes with URL values. Use the pattern href="(http[s]?://.+?)" to identify each URL.
Locate the attribute: The literal text href=" identifies the beginning of an href attribute.
Check the scheme: The expression http[s]? accepts either http or https.
Match the URL body: The non-greedy .+? matches the domain and path but stops when the first closing quote satisfies the pattern.
Capture the URL: The parentheses surround the URL portion, so re.findall() returns that captured text rather than the surrounding href=" and closing quote.
Continue searching: After one match is completed, re.findall() searches for the remaining occurrences and returns the captured URL from each match.
The result is one extracted URL value for each matching href attribute.
Writing the Search Program
import re html = 'href="http://example.com" href="https://example.org/page"' pattern = r'href="(http[s]?://.+?)"' links = re.findall(pattern, html) for link in links: print(link)
http://example.com
https://example.org/pageUsing urllib Data
A complete version of this task can fetch HTML from a URL with urllib and then apply re.findall() to the downloaded content. The source program uses the byte-string pattern b'href="(http[s]?://.*?)"' because urllib returns bytes. The captured URL values are decoded from bytes to strings before they are printed.
Mistakes with URL Patterns
Using .+ instead of .+?
The greedy quantifier tries to match as much as possible and can continue across multiple href values until a later double quote.
Fix:
Use .+? so the match stops at the first closing quote that satisfies the pattern.Leaving out the capture parentheses
Without parentheses around the URL portion, re.findall() returns the entire matching href attribute instead of only the URL.
Fix:
Write href="(http[s]?://.+?)" when the desired result is the URL value.Forgetting the optional s
This pattern does not account for https URLs.
Fix:
Use http[s]? to handle both http and https.Treating complex HTML as a straightforward regex task
The source approach is intended for straightforward extraction; complex or malformed HTML and many attribute edge cases can require more reliable parsing.
Fix:
Consider a dedicated HTML parser library such as BeautifulSoup for complex or malformed HTML.
Practice the Extraction
Write a Python expression that uses re.findall() to extract every http or https URL from an HTML string containing multiple href attributes. Your pattern should return only the URL values, not href=" or the closing quotes.
Hints
- Begin with the literal text href=".
- Use http[s]? to allow both URL schemes.
- Put parentheses around the URL portion.
- Use .+? so each match stops at the first closing quote.
What do you think happens?
What does re.findall() return for the pattern href="(http[s]?://.+?)" when the HTML contains two matching href attributes?
Reveal answer
Answer: One captured URL value for each matching href attribute
The non-greedy quantifier stops each match at the first closing quote, and the parentheses tell re.findall() to return only the URL capture group.
Reliable Use of the Technique
For a straightforward HTML extraction task, build the pattern in layers: anchor it with href=", allow both http and https, make the URL body non-greedy, and place parentheses around the value you want returned. Then use re.findall() to search for every occurrence. If the HTML is complex or malformed, or if many attribute edge cases must be handled reliably, use a dedicated HTML parser library such as BeautifulSoup instead of relying on a regular expression.
Key Takeaways
- The pattern href="(http[s]?://.+?)" targets URL values inside href attributes.
- The optional s supports both http and https URLs.
- The question mark in .+? makes the match non-greedy, so it stops at the first closing quote.
- Parentheses define the capture group returned by re.findall(), allowing the program to extract only the URL.
- This technique suits straightforward HTML extraction; complex or malformed HTML may require a dedicated parser such as BeautifulSoup.