Debugging Web Crawl Results
spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).
From Crawl Database to Evidence
After a web spider finishes crawling, its results are stored in spider.sqlite. That database contains the pages and connectivity information discovered by the spider, but the binary SQLite file is not a convenient format for inspecting those results directly. spdump.py provides a human-readable view by reading spider.sqlite and printing information about the pages and their connections.
The important debugging idea is that spdump.py is not merely showing a list of URLs. Each output line summarizes one page using five values: its incoming-link count, two rank fields, its page identifier, and its URL. Reading those values lets you compare what your spider stored with what you expected the crawl to discover.
Reading the Five Positions
Every output line follows the same order: (incoming_links, old_rank, new_rank, page_id, url). The order matters. A number in the first position is not a page ID, and a number in the fourth position is not an incoming-link count. Interpret each value by its position before drawing conclusions about the crawl.
| Position | Field | Meaning |
|---|---|---|
| 1 | incoming_links | How many other pages in the crawl link to this page |
| 2 | old_rank | The earlier rank value; it is often None on the first run |
| 3 | new_rank | The newer rank value shown by the crawl's PageRank processing |
| 4 | page_id | The page's identifier in the crawled data |
| 5 | url | The page's URL |
Use position and field name together when interpreting each tuple.
Tracing One Tuple
Interpreting an output line
Interpret the source example (5, None, 1.0, 3, url).
Read the first value: The value 5 means that five distinct pages in the crawled network contain a hyperlink pointing to this page.
Read the second value: None in old_rank is consistent with an initial run, before numeric earlier-rank values have appeared after multiple PageRank computations.
Read the third value: The value 1.0 is the new_rank value displayed for this page.
Read the fourth value: The value 3 identifies the page in the crawled data.
Read the fifth value: url identifies the page's URL.
This tuple describes a page with five incoming links, an unavailable earlier rank in this example, a displayed new rank of 1.0, page identifier 3, and the listed URL.
The tuple gives you several ways to cross-check the same page. The URL tells you which page you are inspecting, while page_id identifies that page in the crawled data. The incoming_links value describes how strongly the rest of the crawled network points to it. The rank fields add information about the PageRank processing: old_rank is often None on the first run, while numeric earlier values appear after PageRank has been computed multiple times.
The Output Filter
spdump.py does not display every page discovered by the spider. It includes only pages with at least one incoming link. A page with zero inbound connections is excluded by design, so its absence from the output does not automatically mean that the spider failed to discover it.
This filter focuses the output on pages that participate in the connectivity network. A page may be referenced in a starting URL or discovered during the crawl while still having no other crawled page linking back to it. The source describes such pages as leaf or terminal pages: they can have outgoing links but no incoming links. Because PageRank depends on pages linking to one another, a page with zero incoming links has no way to receive PageRank from other pages in the crawl.
Validating Connectivity
- Identify the URL and page_id for the page you expected to find.
- Read incoming_links first to determine how many other crawled pages point to it.
- Compare that count with your expectation about the page's role in the site structure.
- Inspect old_rank and new_rank together, remembering that old_rank is often None on the first run.
- Treat a missing page carefully: first consider whether it had zero incoming links and was therefore filtered out.
- Use the tuple values to decide whether the crawl's stored connectivity matches the structure you expected.
| Observation | What it can tell you |
|---|---|
| A page has many incoming links | The page is strongly connected within the crawled network and may be a navigation hub, homepage, or popular content page. |
| A page has only one incoming link | The page has a limited observed connection in the crawl; this may be expected or may indicate that the spider did not crawl deeply enough or missed pages. |
| old_rank is None | The output is consistent with an initial run before numeric earlier-rank values have appeared. |
| old_rank and new_rank are both numeric | The output can show the progression of PageRank values across computations. |
| An expected page is absent | The page may have zero inbound connections and therefore be excluded by spdump.py. |
Connectivity counts and rank values answer related but different questions. incoming_links directly reports how many crawled pages link to the page. Rank values show the PageRank processing state. More incoming links tend to support a higher PageRank, but the exact calculation also depends on the PageRank of the pages doing the linking. Therefore, do not treat incoming_links as a complete substitute for rank information.
Mistakes During Inspection
Assuming every crawled page must appear in spdump.py output.
Pages with zero incoming links are filtered out by design.
Fix:
Treat the output as a connectivity-focused view, not a complete page inventory.Reading the tuple without respecting field order.
The fields always appear in the order incoming_links, old_rank, new_rank, page_id, url.
Fix:
Map each value to its position before interpreting it.Treating old_rank being None as proof that the crawl failed.
old_rank is often None on the first run.
Fix:
Check whether numeric values appear after PageRank has been computed multiple times.Assuming a low incoming-link count automatically proves the spider is broken.
The count may reflect the actual crawled network, or it may indicate that the spider did not crawl deeply enough or missed pages.
Fix:
Use the value as a debugging signal and compare it with your expectations about the crawl.
Practice the Diagnosis
A page you expected to be central to the site does not appear in spdump.py output. What should you check first, and what would its absence mean if it had zero incoming links?
Hints
- Start with the output filter rather than assuming the page was never discovered.
- Recall the condition required for a page to appear.
- Distinguish being discovered from having another crawled page link to the page.
A missing central page
Diagnose the missing page using the rules of spdump.py.
Check the inclusion rule: Only pages with at least one incoming link appear in the output.
Interpret the absence: If the page had zero inbound connections, spdump.py would exclude it even if the spider encountered or discovered it.
Choose the next debugging question: Ask whether the crawl actually observed links from other crawled pages to that page. If the expected links were missing, the crawl may not have gone deeply enough or may have missed pages.
The missing page should not immediately be treated as undiscovered or as evidence of a broken spider. First determine whether it had zero incoming links and was filtered out.
Key Takeaways
- spdump.py reads spider.sqlite and presents crawled page information in a human-readable five-element tuple.
- The tuple order is incoming_links, old_rank, new_rank, page_id, url.
- incoming_links is the direct measure of how many other crawled pages link to the page.
- Only pages with at least one incoming link appear, so pages with zero inbound connections are filtered out.
- old_rank is often None on the first run, while numeric earlier values appear after multiple PageRank computations.
- Use the output to compare observed connectivity and rank information with what you expected the spider to discover.
Key Takeaways
- spdump.py turns the crawled information in spider.sqlite into readable page summaries.
- Every displayed tuple uses the fixed order (incoming_links, old_rank, new_rank, page_id, url).
- The incoming-link count is the clearest direct measure of a page's connectivity in the crawl.
- Pages with zero incoming links are intentionally omitted from the output.
- Comparing connectivity, rank fields, identifiers, and URLs helps validate and debug crawl results.