Parsing HTML with BeautifulSoup
Regular expressions provide a concise way to search for and extract patterns in HTML text, such as URLs in href attributes.
From HTML Text to URLs
HTML documents contain structured data embedded in text. A link is commonly represented by an href attribute such as href="http://example.com". If you need to extract several URL values, you could search through the characters manually and build substrings yourself. A regular expression gives you a more concise way to describe the pattern, allowing the regular expression engine to find matching text for you.
Building the Href Pattern
A typical pattern for extracting HTTP and HTTPS links is href="(http[s]?://.+?)". Each part has a specific job. The literal text href=" anchors the search at the beginning of an href attribute. The character sequence http matches the beginning of the protocol. The optional s, written as [s]?, allows either http or https. The characters .+? match the domain and path while using non-greedy matching. The final double quote marks the end of the attribute value.
Greedy and Non-Greedy Matching
The plus quantifier is greedy by default. A greedy match tries to find the largest possible matching string. In contrast, a non-greedy match tries to find the smallest possible matching string. Adding a question mark after the quantifier makes it non-greedy, so .+ becomes .+?.
Capturing the URL Value
The complete href pattern matches more than the URL alone: it includes href=" and the closing quote. Parentheses define a capture group, identifying the part of the match that should be extracted. With parentheses around http[s]?://.+?, re.findall() returns only the URL values. Without those parentheses, the entire href attribute match would be returned.
Selecting one URL from an href
Identify the value extracted from href="http://example.com/page" when the pattern is href="(http[s]?://.+?)".
Match the attribute start: The literal href=" matches the beginning of the attribute.
Match the protocol: http[s]? accepts the HTTP protocol with or without the optional s.
Capture the URL: The parenthesized expression captures http://example.com/page.
Stop at the boundary: The non-greedy .+? stops when it reaches the first closing double quote.
The extracted value is http://example.com/page, not the surrounding href=" and closing quote.
Finding Every Link
Once the pattern is ready, re.findall() searches the HTML text for every occurrence and returns the captured URL values. The loop then processes each extracted value. This is the central flow: provide HTML text, search it with the href pattern, receive a collection of captured URLs, and handle those URLs one at a time.
import re html = ''' <a href="http://example.com">Example</a> <a href="https://example.org/docs">Documentation</a> ''' pattern = r'href="(http[s]?://.+?)"' links = re.findall(pattern, html) for link in links: print(link)
http://example.com
https://example.org/docsMistakes in Link Extraction
Using .+ instead of .+?
The greedy quantifier tries to match the largest possible span and can continue across multiple href values until the last suitable double quote.
Fix:
Use href="(http[s]?://.+?)" so the match stops at the first closing quote.Leaving out the capture parentheses
Without a capture group, the extracted match includes href=" and the closing quote instead of only the URL.
Fix:
Surround the URL portion with parentheses: href="(http[s]?://.+?)".Forgetting the optional s
A pattern containing only http does not account for the https form described in the source.
Fix:
Use http[s]? to allow both http and https.Treating regular expressions as a complete HTML parser
The source describes this approach as suitable for straightforward extraction tasks and recommends a dedicated parser for complex cases and many edge cases.
Fix:
Consider a dedicated HTML parser library such as BeautifulSoup when the HTML structure is complex or malformed.
Pattern Practice
Write a Python regular expression pattern that extracts both http and https URL values from href attributes. Then explain which characters make the URL match non-greedy and which characters define the capture group.
Hints
- Begin with the literal attribute text href=".
- Use http[s]? for both protocol forms.
- Place parentheses around the URL portion.
- Place ? after + to make the match non-greedy.
What do you think happens?
What will re.findall() return for two matching href attributes when the pattern is href="(http[s]?://.+?)"?
Reveal answer
Answer: The two URL values without href=" or the closing quotes
The pattern is non-greedy, so each match stops at the first closing quote, and the parentheses make re.findall() return only the URL portion.
Practical Boundary
For a straightforward HTML extraction task with a known href structure, href="(http[s]?://.+?)" combines the needed ideas: it anchors the search to href=", accepts http and https, captures only the URL, and uses non-greedy matching to stop at the first closing quote. The same technique can be used repeatedly with re.findall() to collect multiple links. When the HTML is complex or malformed, move from this simple pattern-based approach to a dedicated HTML parser such as BeautifulSoup.
Key Takeaways
- The pattern href="(http[s]?://.+?)" targets HTTP and HTTPS URLs inside href attributes.
- The question mark in .+? makes the match non-greedy, so it stops at the first closing quote.
- Parentheses create a capture group, causing re.findall() to return the URL instead of the complete href attribute.
- re.findall() can search HTML text for multiple matching links and return their captured URL values.
- Regular expressions are useful for straightforward extraction, while complex or malformed HTML may require BeautifulSoup or another dedicated HTML parser.