Running Your First Web Spider
spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).
From Crawl Data to Evidence
After a web spider finishes crawling, the next question is not only whether it ran, but what it actually found. The crawl results are stored in spider.sqlite, a binary SQLite database that is not intended to be read directly as plain text. spdump.py provides a human-readable view of the stored page relationships, incoming-link counts, and PageRank values. It is therefore useful both for understanding the discovered network and for checking whether the crawl behaved as expected.
Reading the Five Positions
Every displayed line follows one fixed structure: (incoming_links, old_rank, new_rank, page_id, url). Each tuple describes one page that passed spdump.py's output filter. The position matters: the same value can mean something completely different if you read it from the wrong position.
| Position | Field | Meaning |
|---|---|---|
| 1 | incoming_links | How many other pages in the crawl link to this page |
| 2 | old_rank | The earlier PageRank value; often None on the first run |
| 3 | new_rank | The newer PageRank value shown by the tool |
| 4 | page_id | The stored identifier for the page |
| 5 | url | The page's URL |
Read the tuple from left to right; do not reorder the fields when interpreting a line.
Interpreting one tuple
Interpret the source example (5, None, 1.0, 3, url).
First position: The value 5 means that five distinct pages in the crawled network contain a hyperlink pointing to this page.
Second position: None in old_rank is common on the first run, before earlier numeric PageRank values have been accumulated.
Third position: The value 1.0 is the new_rank value displayed for this page.
Final positions: The value 3 is the page_id, and url identifies the page's URL.
This tuple describes a page with five incoming links, no earlier rank value available, a displayed new rank of 1.0, page ID 3, and the shown URL.
Connectivity Through Incoming Links
The first tuple element is the most direct measure of page connectivity. It counts how many other pages in the crawl link to the page represented by that tuple. A larger incoming-link count means that more pages in the crawled network point toward that page. Pages with many incoming links may be navigation hubs, homepages, or popular content pages, although the count alone does not determine the exact PageRank.
The Output Filter
spdump.py does not print every page discovered by the spider. It displays only pages with at least one incoming link. A page with zero inbound connections is excluded by design, so absence from the output does not by itself prove that the spider failed to discover that page.
This filter keeps attention on pages that participate in the connectivity network. The source describes pages with no incoming links as leaf or terminal pages: they may have outgoing links, but no other page in the crawled network links back to them. Because they receive no links from other pages, they cannot receive PageRank from those pages.
Assuming that every discovered page must appear in spdump.py output.
The tool filters out pages with no inbound connections.
Fix:
Interpret absence together with the output rule: only pages with at least one incoming link are displayed.Treating a missing page as proof that the spider did not find it.
The page may have been discovered but excluded because no crawled page links to it.
Fix:
Distinguish discovery from display. spdump.py shows a filtered connectivity summary, not an unfiltered list of every page.
Validation and Debugging
Use the tuples as evidence when checking a crawl. Compare the URLs shown with the pages you expected to discover. Inspect incoming-link counts to see whether an expected hub has meaningful connectivity. Check page IDs and rank fields for values that do not fit your understanding of the stored crawl. Unexpected links, surprisingly small link counts, missing URLs, or unexpected rank values can point to a crawl that did not go as deeply as expected, missed pages, or stored relationships different from the ones you intended to examine.
What do you think happens?
A page was discovered during the crawl, but no crawled page links to it. Will it appear in spdump.py output?
Reveal answer
Answer: No, it is filtered out.
spdump.py displays only pages with at least one incoming link. A discovered page with zero inbound connections is excluded by design.
- Verify that an expected URL is present when it should participate in the displayed connectivity network.
- Read the first tuple element to check whether the observed incoming-link count matches the crawl structure you expected.
- Use page_id and url together to identify which stored page a tuple describes.
- Treat old_rank being None on the first run as an expected condition rather than automatically as an error.
- Remember that missing output may result from the zero-incoming-link filter.
Practice Reading Tuples
Suppose a displayed tuple has the form (2, None, 0.5, 7, url). Explain what each of the five positions tells you, then state whether this page passed the spdump.py output filter and why.
Hints
- Start with the first position because it determines whether the page has incoming connectivity.
- Use the tuple order: incoming_links, old_rank, new_rank, page_id, url.
- A page must have at least one incoming link to appear.
Practice answer
Interpret (2, None, 0.5, 7, url).
Connectivity: The page has two incoming links, meaning two other pages in the crawl link to it.
Ranks: old_rank is None, while new_rank is 0.5. None is consistent with an initial run before earlier numeric rank values are available.
Identity: The page_id is 7 and the final field is the page URL.
Filter decision: The page appears because its incoming-link count is two, which is at least one.
The tuple represents a displayed page with two incoming links, an unavailable earlier rank, a new rank of 0.5, page ID 7, and the shown URL.
Key Takeaways
- spdump.py reads spider.sqlite and presents a human-readable connectivity summary.
- Each output line follows the order (incoming_links, old_rank, new_rank, page_id, url).
- incoming_links counts how many other crawled pages link to the page represented by the tuple.
- Only pages with at least one incoming link appear; pages with zero inbound connections are filtered out.
- Comparing URLs, link counts, IDs, and rank fields with your expectations helps validate and debug a crawl.