Concepts / Introduction to urllib and HTTP Requests

Introduction to urllib and HTTP Requests

BeautifulSoup transforms raw HTML strings into searchable tree-structured objects that mirror the document's logical hierarchy.

  • Programming

From Web Response to Searchable Data

A web page initially arrives as raw HTML: a long string of characters. That string may describe headings, links, and other document elements, but it is unstructured from a programmatic perspective. The central workflow is to fetch the HTML with urllib, pass it to BeautifulSoup with a parser specification, and then search the resulting object for the elements you need.

HTTP requestHTML responseurllibWeb serverRaw HTMLlong text string
How does a request travel from urllib to a web server and return HTML data for parsing?

Building the BeautifulSoup Tree

BeautifulSoup transforms a raw HTML string into a searchable tree-structured object. The tree mirrors the document's logical hierarchy, so the document is no longer treated as one flat sequence of characters. Instead, its elements can be navigated and searched as tags. The parser='html.parser' argument specifies Python's built-in HTML parser.

python
containscontainscontainsHTML documenthtmlbodyaHome
How does a raw HTML string become a nested, searchable tree of document elements?

In this generated example, the HTML string is the input and soup represents the complete parsed document. The important change is not that the text became shorter; it became organized into relationships that BeautifulSoup can search.

Selecting Anchor Tags

After parsing, call the soup object with a tag name. The soup(tag_name) method retrieves all tags of that type. Therefore, soup('a') returns a list containing every anchor tag in the document. This search happens after parsing; urllib supplies the raw HTML, while BeautifulSoup supplies the structure and search operation.

anchors = soup('a')

search with soup('a')returns all matchesBeautifulSoupobjectparsed documentarequested tagAnchor tagsmatching list
How does BeautifulSoup search the parsed HTML tree and select only the requested tags?

Reading href Safely

An anchor tag may contain an href attribute, but real-world HTML can also contain an anchor without that attribute. Use tag.get('href', None) to request the attribute safely. If href is absent, get() returns the default value None instead of raising an error.

python
get('href', None)get('href', None)aHome linkhome.htmlhref valueaincomplete linkNonedefault value
How do anchor elements in parsed HTML map to their individual href attribute values?
Access patternWhen href existsWhen href is missing
tag.get('href', None)Returns the href valueReturns None
tag['href']Returns the href valueRaises KeyError

Certificate Verification Failures

Fetching a page can encounter an SSL certificate error. One documented handling approach is to create and configure an SSL context that bypasses certificate verification. This changes the request flow so the fetch can proceed in situations where verification would otherwise fail.

beginsfailsconfigure handlingrequest proceedsWeb requestCertificateverificationSSL certificate errorConfigured SSLcontextverification bypassedHTML responseavailable for parsing
What happens in the request flow when SSL certificate verification fails, and where can handling change the outcome?

Installing the Parsing Library

BeautifulSoup must be installed before a Python program can use it to parse HTML. Installation makes the library available to the program; parsing begins afterward, when the fetched raw HTML is passed to BeautifulSoup with a parser specification. The source workflow assumes both stages: first make the library available, then fetch and parse the page.

  1. Install BeautifulSoup using the package-management process for your Python environment.
  2. Fetch the web page so urllib provides raw HTML text.
  3. Pass the raw HTML to BeautifulSoup with parser='html.parser'.
  4. Search the parsed object with soup('a') or another tag name.
  5. Read optional attributes with tag.get(attribute_name, default_value).

Mistakes That Break Extraction

  • Treating the fetched HTML as if it were already a searchable document

    urllib provides raw HTML as a long string. BeautifulSoup is the step that transforms it into a searchable tree-structured object.

    Fix: Pass the raw HTML to BeautifulSoup with parser='html.parser' before searching.

  • Searching for anchor tags before parsing

    The soup(tag_name) search pattern applies to the parsed BeautifulSoup object.

    Fix: Create the parsed object first, then call soup('a').

  • Assuming every anchor has an href attribute

    Bracket notation raises KeyError when the attribute is missing.

    Fix: Use tag.get('href', None) so a missing attribute produces the default None.

  • Treating SSL verification bypass as a production default

    The source identifies bypassing verification as useful for educational or testing environments, while production code should verify certificates.

    Fix: Use a configured bypass only when necessary in the relevant environment, and verify certificates in production.

Practice the Workflow

EASY

Suppose a fetched page has been parsed into soup. Write the extraction logic that finds every anchor tag and obtains each tag's href value without failing when href is absent.

Hints
  • Use soup with the tag name a.
  • Process the returned tags one at a time.
  • Use get() with href and None as its arguments.

What do you think happens?

What value should tag.get('href', None) provide when the current anchor tag has no href attribute?

  • The empty HTML document
  • None
  • A KeyError
  • Every anchor tag in the page
Reveal answer

Answer: None

The second argument to get() is the default value. It is returned when the requested attribute is missing, whereas bracket notation would raise KeyError.

Workflow Summary

  1. urllib fetches a page and provides raw HTML as text.
  2. BeautifulSoup transforms that text into a searchable tree-structured object.
  3. soup('a') returns a list of all anchor tags in the parsed document.
  4. tag.get('href', None) safely retrieves href and returns None when it is missing.
  5. SSL errors can be handled with a configured SSL context, but production code should verify certificates.

Key Takeaways

  • Fetching and parsing are separate stages: urllib supplies raw HTML, and BeautifulSoup organizes it into a searchable tree.
  • Use soup('a') to retrieve all anchor tags from the parsed document.
  • Use tag.get('href', None) to handle missing href attributes without a KeyError.
  • An SSL context can bypass verification for some educational or testing situations, while production code should verify certificates.
  • BeautifulSoup must be installed before the parsing workflow can be used.