Concepts / Parsing HTML with BeautifulSoup

Parsing HTML with BeautifulSoup

Regular expressions provide a concise way to search for and extract patterns in HTML text, such as URLs in href attributes.

  • Programming

From HTML Text to URLs

HTML documents contain structured data embedded in text. A link is commonly represented by an href attribute such as href="http://example.com". If you need to extract several URL values, you could search through the characters manually and build substrings yourself. A regular expression gives you a more concise way to describe the pattern, allowing the regular expression engine to find matching text for you.

Building the Href Pattern

A typical pattern for extracting HTTP and HTTPS links is href="(http[s]?://.+?)". Each part has a specific job. The literal text href=" anchors the search at the beginning of an href attribute. The character sequence http matches the beginning of the protocol. The optional s, written as [s]?, allows either http or https. The characters .+? match the domain and path while using non-greedy matching. The final double quote marks the end of the attribute value.

href="attribute anchorhttpprotocol text[s]?optional s://separator(.+?)captured URL"closing boundary
How do the literal characters, optional character, wildcard, quantifier, and capture group work together?
python

Greedy and Non-Greedy Matching

The plus quantifier is greedy by default. A greedy match tries to find the largest possible matching string. In contrast, a non-greedy match tries to find the smallest possible matching string. Adding a question mark after the quantifier makes it non-greedy, so .+ becomes .+?.

greedy searchstop at first quoteTwo href attributesmultiple closing quotesTwo href attributessame HTML textGreedy .+larger spanNon-greedy .+?first URL span
What substring does each pattern match when HTML contains multiple href attributes?

Capturing the URL Value

The complete href pattern matches more than the URL alone: it includes href=" and the closing quote. Parentheses define a capture group, identifying the part of the match that should be extracted. With parentheses around http[s]?://.+?, re.findall() returns only the URL values. Without those parentheses, the entire href attribute match would be returned.

parentheses selecthref="http://example.com"entire matchhttp://example.comcapture group result
What is the difference between the entire href match and the smaller URL returned from its capture group?

Selecting one URL from an href

Identify the value extracted from href="http://example.com/page" when the pattern is href="(http[s]?://.+?)".

Match the attribute start: The literal href=" matches the beginning of the attribute.

Match the protocol: http[s]? accepts the HTTP protocol with or without the optional s.

Capture the URL: The parenthesized expression captures http://example.com/page.

Stop at the boundary: The non-greedy .+? stops when it reaches the first closing double quote.

The extracted value is http://example.com/page, not the surrounding href=" and closing quote.

Finding Every Link

Once the pattern is ready, re.findall() searches the HTML text for every occurrence and returns the captured URL values. The loop then processes each extracted value. This is the central flow: provide HTML text, search it with the href pattern, receive a collection of captured URLs, and handle those URLs one at a time.

import re html = ''' <a href="http://example.com">Example</a> <a href="https://example.org/docs">Documentation</a> ''' pattern = r'href="(http[s]?://.+?)"' links = re.findall(pattern, html) for link in links: print(link)

Output
http://example.com
https://example.org/docs
find firstcontinue searchappend resultHTML texttwo href attributeshttp://example.comfirst captured valuehttps://example.org/docssecond captured valuelinkscollected URL values
How does the program move through the HTML text, find each matching link, and collect the URL values?

Mistakes in Link Extraction

  • Using .+ instead of .+?

    The greedy quantifier tries to match the largest possible span and can continue across multiple href values until the last suitable double quote.

    Fix: Use href="(http[s]?://.+?)" so the match stops at the first closing quote.

  • Leaving out the capture parentheses

    Without a capture group, the extracted match includes href=" and the closing quote instead of only the URL.

    Fix: Surround the URL portion with parentheses: href="(http[s]?://.+?)".

  • Forgetting the optional s

    A pattern containing only http does not account for the https form described in the source.

    Fix: Use http[s]? to allow both http and https.

  • Treating regular expressions as a complete HTML parser

    The source describes this approach as suitable for straightforward extraction tasks and recommends a dedicated parser for complex cases and many edge cases.

    Fix: Consider a dedicated HTML parser library such as BeautifulSoup when the HTML structure is complex or malformed.

Pattern Practice

MEDIUM

Write a Python regular expression pattern that extracts both http and https URL values from href attributes. Then explain which characters make the URL match non-greedy and which characters define the capture group.

Hints
  • Begin with the literal attribute text href=".
  • Use http[s]? for both protocol forms.
  • Place parentheses around the URL portion.
  • Place ? after + to make the match non-greedy.

What do you think happens?

What will re.findall() return for two matching href attributes when the pattern is href="(http[s]?://.+?)"?

  • The complete HTML document
  • The two URL values without href=" or the closing quotes
  • One string containing both href attributes
  • Only the first URL
Reveal answer

Answer: The two URL values without href=" or the closing quotes

The pattern is non-greedy, so each match stops at the first closing quote, and the parentheses make re.findall() return only the URL portion.

Practical Boundary

For a straightforward HTML extraction task with a known href structure, href="(http[s]?://.+?)" combines the needed ideas: it anchors the search to href=", accepts http and https, captures only the URL, and uses non-greedy matching to stop at the first closing quote. The same technique can be used repeatedly with re.findall() to collect multiple links. When the HTML is complex or malformed, move from this simple pattern-based approach to a dedicated HTML parser such as BeautifulSoup.

Key Takeaways

  • The pattern href="(http[s]?://.+?)" targets HTTP and HTTPS URLs inside href attributes.
  • The question mark in .+? makes the match non-greedy, so it stops at the first closing quote.
  • Parentheses create a capture group, causing re.findall() to return the URL instead of the complete href attribute.
  • re.findall() can search HTML text for multiple matching links and return their captured URL values.
  • Regular expressions are useful for straightforward extraction, while complex or malformed HTML may require BeautifulSoup or another dedicated HTML parser.