Concepts / Working with URLs and urllib

Working with URLs and urllib

Regular expressions provide a concise way to search for and extract patterns in HTML text, such as URLs in href attributes.

  • Programming

From HTML Text to URLs

HTML contains structured data embedded in text. A link commonly stores its destination in an href attribute, such as href="http://example.com". If you need every link in a page, you could search through characters manually and build substrings, but a regular expression lets you describe the pattern and ask the regex engine to find matching text.

The extraction task has two layers. First, the pattern must recognize the href attribute and its URL. Second, the pattern must identify which part of that match should be returned. The pattern href="(http[s]?://.+?)" handles both layers: the literal href=" anchors the search, while parentheses surround the URL that should be extracted.

thencontainsstops beforehref=attribute anchor"opening quote(http[s]?://.+?)captured URL"closing quote
How does the regular expression move through an href attribute to identify and isolate the URL value?

Reading the Pattern

In href="(http[s]?://.+?)", href=" is literal text that anchors the match to the beginning of an href attribute. The expression [s]? allows either http or https. The dot and plus, .+, match the domain and path, and the question mark after the plus changes that quantifier from greedy to non-greedy.

Pattern partPurpose
href="Anchors the search to the beginning of an href attribute
httpRequires the URL scheme to begin with http
[s]?Allows the optional s in https
.+?Matches the domain and path non-greedily
( ... )Defines the part returned as a capture group
"Marks the closing quote of the attribute

Parts of the URL-extraction pattern

The pattern must stop at the closing quote belonging to the current href attribute. That stopping behavior matters when several links occur in the same HTML text. A greedy quantifier tries to match the largest possible string, while a non-greedy quantifier tries to match the smallest possible string that still satisfies the pattern.

matches as much as possiblematches as little as possible.+greedyfirst URL throughlast quotepotentially multiple hrefvalues.+?non-greedyone URLstops at first quote
What substring does the pattern match when it is greedy versus non-greedy, and why does the non-greedy version stop at the correct closing quote?

Tracing Two Links

Extracting Two URL Values

Suppose HTML contains two href attributes with URL values. Use the pattern href="(http[s]?://.+?)" to identify each URL.

Locate the attribute: The literal text href=" identifies the beginning of an href attribute.

Check the scheme: The expression http[s]? accepts either http or https.

Match the URL body: The non-greedy .+? matches the domain and path but stops when the first closing quote satisfies the pattern.

Capture the URL: The parentheses surround the URL portion, so re.findall() returns that captured text rather than the surrounding href=" and closing quote.

Continue searching: After one match is completed, re.findall() searches for the remaining occurrences and returns the captured URL from each match.

The result is one extracted URL value for each matching href attribute.

containsparentheses extracthref attributehref="URL"URLhttp or https valuefull matchhref="URL"capture groupURL only
What contains what: how is the URL nested inside the href attribute, and how does the match represent those layers?
parentheses selecthref="http://example.com"full matchhttp://example.comcaptured URL
Which part of the matched HTML is captured by the parentheses, and how does that captured value differ from the full match?

Writing the Search Program

import re html = 'href="http://example.com" href="https://example.org/page"' pattern = r'href="(http[s]?://.+?)"' links = re.findall(pattern, html) for link in links: print(link)

Output
http://example.com
https://example.org/page
searchfindsfindscollectscollectsHTML textmultiple href attributesre.findall()search patternURL 1captured valueURL 2captured valuelinkscollected results
How does the program scan the HTML, find each matching href, extract its URL, and collect the results?

Using urllib Data

A complete version of this task can fetch HTML from a URL with urllib and then apply re.findall() to the downloaded content. The source program uses the byte-string pattern b'href="(http[s]?://.*?)"' because urllib returns bytes. The captured URL values are decoded from bytes to strings before they are printed.

python

Mistakes with URL Patterns

  • Using .+ instead of .+?

    The greedy quantifier tries to match as much as possible and can continue across multiple href values until a later double quote.

    Fix: Use .+? so the match stops at the first closing quote that satisfies the pattern.

  • Leaving out the capture parentheses

    Without parentheses around the URL portion, re.findall() returns the entire matching href attribute instead of only the URL.

    Fix: Write href="(http[s]?://.+?)" when the desired result is the URL value.

  • Forgetting the optional s

    This pattern does not account for https URLs.

    Fix: Use http[s]? to handle both http and https.

  • Treating complex HTML as a straightforward regex task

    The source approach is intended for straightforward extraction; complex or malformed HTML and many attribute edge cases can require more reliable parsing.

    Fix: Consider a dedicated HTML parser library such as BeautifulSoup for complex or malformed HTML.

Practice the Extraction

MEDIUM

Write a Python expression that uses re.findall() to extract every http or https URL from an HTML string containing multiple href attributes. Your pattern should return only the URL values, not href=" or the closing quotes.

Hints
  • Begin with the literal text href=".
  • Use http[s]? to allow both URL schemes.
  • Put parentheses around the URL portion.
  • Use .+? so each match stops at the first closing quote.

What do you think happens?

What does re.findall() return for the pattern href="(http[s]?://.+?)" when the HTML contains two matching href attributes?

  • One string containing both href attributes
  • The complete href attributes including href=" and the closing quotes
  • One captured URL value for each matching href attribute
Reveal answer

Answer: One captured URL value for each matching href attribute

The non-greedy quantifier stops each match at the first closing quote, and the parentheses tell re.findall() to return only the URL capture group.

Reliable Use of the Technique

For a straightforward HTML extraction task, build the pattern in layers: anchor it with href=", allow both http and https, make the URL body non-greedy, and place parentheses around the value you want returned. Then use re.findall() to search for every occurrence. If the HTML is complex or malformed, or if many attribute edge cases must be handled reliably, use a dedicated HTML parser library such as BeautifulSoup instead of relying on a regular expression.

Key Takeaways

  • The pattern href="(http[s]?://.+?)" targets URL values inside href attributes.
  • The optional s supports both http and https URLs.
  • The question mark in .+? makes the match non-greedy, so it stops at the first closing quote.
  • Parentheses define the capture group returned by re.findall(), allowing the program to extract only the URL.
  • This technique suits straightforward HTML extraction; complex or malformed HTML may require a dedicated parser such as BeautifulSoup.