Concepts / Handling Errors in Web Scraping

Handling Errors in Web Scraping

BeautifulSoup's soup('a') method finds all anchor tags in parsed HTML; tag.get('href', None) safely extracts the href attribute from each tag.

  • Programming

From Page to Link List

A web page displays clickable links, but a scraper sees those links as HTML anchor tags. An anchor tag stores its destination in an href attribute. BeautifulSoup can parse downloaded HTML, find every anchor tag, and retrieve each href value. When something goes wrong, the reliable approach is to trace the data from downloaded bytes to the final printed values instead of guessing which step failed.

parsesoup('a')get('href', None)printRaw HTML bytesBeautifulSoup treeAnchor tagshref valuesPrinted link list
How does data move from raw HTML through BeautifulSoup parsing and anchor-tag selection to the final list of extracted URLs?

Tracing the Extraction Steps

  1. urllib.request.urlopen() opens the address over the network.
  2. read() obtains the downloaded HTML as bytes.
  3. BeautifulSoup parses those bytes and builds an internal tree representing the HTML document.
  4. soup('a') searches the parsed tree and returns all anchor tag elements.
  5. tag.get('href', None) retrieves the href value from each anchor tag, or returns None when that attribute is absent.
  6. print() sends each extracted value to the console.
getgetgetAnchor 1<ahref="https://example.com">https://example.comhrefAnchor 2<ahref="tutorial/index.html">tutorial/index.htmlhrefAnchor 3<a>Nonemissing href
Which href value belongs to each anchor tag, and what happens when an anchor has no href attribute?

The expression soup('a') searches the entire parsed document tree, not just one particular part of the page. It returns a list of all matching anchor tag elements. Each returned tag can contain attributes such as href. Calling tag.get('href', None) asks for the href attribute safely: an existing attribute produces its stored value, while a missing attribute produces None.

python
Output
https://example.com
None
containscontainscontainscontainsmatchmatchHTML documentbodyahttps://example.comatutorial/index.htmlAnchor tag listsection
What elements does soup('a') find in the parsed document, and how are multiple matching anchor tags returned?

Certificate Verification

When urllib.request.urlopen() opens a URL, Python checks the website's SSL/TLS certificate. The check is intended to verify that the connection is secure and that the connection is with the real server rather than an imposter. In some websites or network environments, this check can fail even when the website is legitimate.

python
HTTPS requestpassesfailsexplicit configurationcontext parameterurlopen()HTML bytesverification succeedsCertificate checkCertificate errorverification failsSSL contextcheck_hostname=False,CERT_NONEHTML bytesconfigured context passedto urlopen()
What changes in the request flow when certificate verification succeeds, fails, or is explicitly configured?

Debugging the Data Flow

A link extractor can fail at several different points. The HTML might not have been downloaded or parsed correctly. The search might find fewer anchor tags than expected. Or the anchor tags might exist without href attributes. Add observations at each stage so the first unexpected value identifies the part of the pipeline that needs investigation.

html = response.read() print("downloaded bytes:", len(html)) soup = BeautifulSoup(html, "html.parser") print("soup created:", soup is not None) tags = soup('a') print("anchor tags:", len(tags)) for tag in tags: href = tag.get('href', None) print("href:", href)

ObservationWhat it indicatesWhere to investigate
len(html) is 0The download produced no HTML bytesThe download step
len(tags) is 0No anchor tags were foundThe page content or parsing step
href is None for every tagAnchor tags exist without href attributesThe HTML attributes
Some href values appearThe extraction reached the attribute valuesThe final output and expected values

Use observations at each stage to locate the first divergence in link extraction.

returnsreturnshref presenttag.get('href', None)https://example.comattribute valuehref missingtag.get('href', None)Nonedefault value
How does get('href', None) return a URL when href exists and None when it is missing?

Mistakes to Avoid

  • Accessing href directly instead of using get()

    An anchor tag can exist without an href attribute.

    Fix: Use tag.get('href', None) so a missing attribute produces None.

  • Forgetting to pass the SSL context

    The configured context is not used unless it is passed to urlopen().

    Fix: Pass the context through the context parameter when opening the URL.

  • Debugging only the final printed list

    The failure may have occurred during download, parsing, or tag selection rather than attribute retrieval.

    Fix: Trace len(html), soup creation, len(tags), and each href value.

  • Assuming every href is a complete web address

    The program prints the exact href stored in each anchor tag, including relative paths and fragments.

    Fix: Interpret the output as raw attribute values before deciding how to process them.

Practice the Trace

EASY

Suppose a parsed document contains three anchor tags. The first has href="/guide", the second has no href attribute, and the third has href="https://example.com/next". Predict the three values printed by a loop that uses tag.get('href', None). Then identify which debug value would tell you how many anchor tags were found.

Hints
  • The loop prints the stored href value for each tag in order.
  • A missing href attribute produces None.
  • The number of matching anchor tags is obtained with len(tags).

What do you think happens?

What does the loop print for an anchor tag that has no href attribute?

  • An empty string
  • None
  • The visible anchor text
  • A certificate error
Reveal answer

Answer: None

tag.get('href', None) returns the href value when the attribute exists and returns None when the attribute is absent.

Reliable Extraction

  1. The extraction pipeline moves from downloaded HTML bytes to a BeautifulSoup tree, then to anchor tags, href values, and printed output.
  2. soup('a') searches the complete parsed document tree and returns all anchor tag elements it finds.
  3. tag.get('href', None) safely retrieves an href value and returns None when the attribute is missing.
  4. Certificate verification can be explicitly disabled with an SSL context, but doing so has security implications.
  5. Debugging is most effective when you inspect download size, soup creation, tag count, and extracted href values in sequence.

Key Takeaways

  • BeautifulSoup turns downloaded HTML bytes into a searchable document tree.
  • soup('a') returns all anchor tags in that parsed tree.
  • tag.get('href', None) returns an href value or None when href is absent.
  • SSL/TLS certificate handling affects whether urlopen() can complete the request.
  • Tracing each intermediate value reveals where link extraction diverges from expectations.