Concepts / Debugging Web Crawl Results

Debugging Web Crawl Results

spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).

  • Programming

From Crawl Database to Evidence

After a web spider finishes crawling, its results are stored in spider.sqlite. That database contains the pages and connectivity information discovered by the spider, but the binary SQLite file is not a convenient format for inspecting those results directly. spdump.py provides a human-readable view by reading spider.sqlite and printing information about the pages and their connections.

The important debugging idea is that spdump.py is not merely showing a list of URLs. Each output line summarizes one page using five values: its incoming-link count, two rank fields, its page identifier, and its URL. Reading those values lets you compare what your spider stored with what you expected the crawl to discover.

recordsdisplaysspider.sqlitecrawled recordsspdump.pyreads databaseFive-element tuplepage summary
How does crawled database information move through spdump.py and become a five-element output line?

Reading the Five Positions

Every output line follows the same order: (incoming_links, old_rank, new_rank, page_id, url). The order matters. A number in the first position is not a page ID, and a number in the fourth position is not an incoming-link count. Interpret each value by its position before drawing conclusions about the crawl.

Position 1incoming_linksPosition 2old_rankPosition 3new_rankPosition 4page_idPosition 5url
What does each position in the tuple represent, and how does each position map to its field name?
PositionFieldMeaning
1incoming_linksHow many other pages in the crawl link to this page
2old_rankThe earlier rank value; it is often None on the first run
3new_rankThe newer rank value shown by the crawl's PageRank processing
4page_idThe page's identifier in the crawled data
5urlThe page's URL

Use position and field name together when interpreting each tuple.

Tracing One Tuple

Interpreting an output line

Interpret the source example (5, None, 1.0, 3, url).

Read the first value: The value 5 means that five distinct pages in the crawled network contain a hyperlink pointing to this page.

Read the second value: None in old_rank is consistent with an initial run, before numeric earlier-rank values have appeared after multiple PageRank computations.

Read the third value: The value 1.0 is the new_rank value displayed for this page.

Read the fourth value: The value 3 identifies the page in the crawled data.

Read the fifth value: url identifies the page's URL.

This tuple describes a page with five incoming links, an unavailable earlier rank in this example, a displayed new rank of 1.0, page identifier 3, and the listed URL.

The tuple gives you several ways to cross-check the same page. The URL tells you which page you are inspecting, while page_id identifies that page in the crawled data. The incoming_links value describes how strongly the rest of the crawled network points to it. The rank fields add information about the PageRank processing: old_rank is often None on the first run, while numeric earlier values appear after PageRank has been computed multiple times.

hyperlinkhyperlinkhyperlinkPage Alinks to targetTarget pageincoming_links = 3Page Blinks to targetPage Clinks to target
How can an incoming-link count be traced back to the pages that point to the target page?

The Output Filter

spdump.py does not display every page discovered by the spider. It includes only pages with at least one incoming link. A page with zero inbound connections is excluded by design, so its absence from the output does not automatically mean that the spider failed to discover it.

This filter focuses the output on pages that participate in the connectivity network. A page may be referenced in a starting URL or discovered during the crawl while still having no other crawled page linking back to it. The source describes such pages as leaf or terminal pages: they can have outgoing links but no incoming links. Because PageRank depends on pages linking to one another, a page with zero incoming links has no way to receive PageRank from other pages in the crawl.

inspectyesnoCrawled pagerecord in spider.sqliteIncoming linksat least one?Tuple outputincludedZero inbound linksfiltered out
What happens to a crawled page when spdump.py decides whether to include it in the displayed output?

Validating Connectivity

  1. Identify the URL and page_id for the page you expected to find.
  2. Read incoming_links first to determine how many other crawled pages point to it.
  3. Compare that count with your expectation about the page's role in the site structure.
  4. Inspect old_rank and new_rank together, remembering that old_rank is often None on the first run.
  5. Treat a missing page carefully: first consider whether it had zero incoming links and was therefore filtered out.
  6. Use the tuple values to decide whether the crawl's stored connectivity matches the structure you expected.
ObservationWhat it can tell you
A page has many incoming linksThe page is strongly connected within the crawled network and may be a navigation hub, homepage, or popular content page.
A page has only one incoming linkThe page has a limited observed connection in the crawl; this may be expected or may indicate that the spider did not crawl deeply enough or missed pages.
old_rank is NoneThe output is consistent with an initial run before numeric earlier-rank values have appeared.
old_rank and new_rank are both numericThe output can show the progression of PageRank values across computations.
An expected page is absentThe page may have zero inbound connections and therefore be excluded by spdump.py.
PageRank progressionold_rankearlier value or Nonenew_ranknewer value
What changes between old_rank and new_rank, and how can the comparison help you understand PageRank processing?

Connectivity counts and rank values answer related but different questions. incoming_links directly reports how many crawled pages link to the page. Rank values show the PageRank processing state. More incoming links tend to support a higher PageRank, but the exact calculation also depends on the PageRank of the pages doing the linking. Therefore, do not treat incoming_links as a complete substitute for rank information.

Mistakes During Inspection

  • Assuming every crawled page must appear in spdump.py output.

    Pages with zero incoming links are filtered out by design.

    Fix: Treat the output as a connectivity-focused view, not a complete page inventory.

  • Reading the tuple without respecting field order.

    The fields always appear in the order incoming_links, old_rank, new_rank, page_id, url.

    Fix: Map each value to its position before interpreting it.

  • Treating old_rank being None as proof that the crawl failed.

    old_rank is often None on the first run.

    Fix: Check whether numeric values appear after PageRank has been computed multiple times.

  • Assuming a low incoming-link count automatically proves the spider is broken.

    The count may reflect the actual crawled network, or it may indicate that the spider did not crawl deeply enough or missed pages.

    Fix: Use the value as a debugging signal and compare it with your expectations about the crawl.

Practice the Diagnosis

MEDIUM

A page you expected to be central to the site does not appear in spdump.py output. What should you check first, and what would its absence mean if it had zero incoming links?

Hints
  • Start with the output filter rather than assuming the page was never discovered.
  • Recall the condition required for a page to appear.
  • Distinguish being discovered from having another crawled page link to the page.

A missing central page

Diagnose the missing page using the rules of spdump.py.

Check the inclusion rule: Only pages with at least one incoming link appear in the output.

Interpret the absence: If the page had zero inbound connections, spdump.py would exclude it even if the spider encountered or discovered it.

Choose the next debugging question: Ask whether the crawl actually observed links from other crawled pages to that page. If the expected links were missing, the crawl may not have gone deeply enough or may have missed pages.

The missing page should not immediately be treated as undiscovered or as evidence of a broken spider. First determine whether it had zero incoming links and was filtered out.

Key Takeaways

  1. spdump.py reads spider.sqlite and presents crawled page information in a human-readable five-element tuple.
  2. The tuple order is incoming_links, old_rank, new_rank, page_id, url.
  3. incoming_links is the direct measure of how many other crawled pages link to the page.
  4. Only pages with at least one incoming link appear, so pages with zero inbound connections are filtered out.
  5. old_rank is often None on the first run, while numeric earlier values appear after multiple PageRank computations.
  6. Use the output to compare observed connectivity and rank information with what you expected the spider to discover.

Key Takeaways

  • spdump.py turns the crawled information in spider.sqlite into readable page summaries.
  • Every displayed tuple uses the fixed order (incoming_links, old_rank, new_rank, page_id, url).
  • The incoming-link count is the clearest direct measure of a page's connectivity in the crawl.
  • Pages with zero incoming links are intentionally omitted from the output.
  • Comparing connectivity, rank fields, identifiers, and URLs helps validate and debug crawl results.