Concepts / Running Your First Web Spider

Running Your First Web Spider

spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).

  • Programming

From Crawl Data to Evidence

After a web spider finishes crawling, the next question is not only whether it ran, but what it actually found. The crawl results are stored in spider.sqlite, a binary SQLite database that is not intended to be read directly as plain text. spdump.py provides a human-readable view of the stored page relationships, incoming-link counts, and PageRank values. It is therefore useful both for understanding the discovered network and for checking whether the crawl behaved as expected.

readsdisplaysspider.sqlitestored crawl dataspdump.pyreads databaseFive-element tuplespage summary
How does spdump.py transform stored crawl data into a readable summary?

Reading the Five Positions

Every displayed line follows one fixed structure: (incoming_links, old_rank, new_rank, page_id, url). Each tuple describes one page that passed spdump.py's output filter. The position matters: the same value can mean something completely different if you read it from the wrong position.

position 1 to 2position 2 to 3position 3 to 4position 4 to 5incoming_linksnumber of linking pagesold_rankprevious rank valuenew_rankcurrent rank valuepage_idpage identifierurlpage address
How does each tuple position map to the page's connectivity, rank values, identifier, and URL?
PositionFieldMeaning
1incoming_linksHow many other pages in the crawl link to this page
2old_rankThe earlier PageRank value; often None on the first run
3new_rankThe newer PageRank value shown by the tool
4page_idThe stored identifier for the page
5urlThe page's URL

Read the tuple from left to right; do not reorder the fields when interpreting a line.

Interpreting one tuple

Interpret the source example (5, None, 1.0, 3, url).

First position: The value 5 means that five distinct pages in the crawled network contain a hyperlink pointing to this page.

Second position: None in old_rank is common on the first run, before earlier numeric PageRank values have been accumulated.

Third position: The value 1.0 is the new_rank value displayed for this page.

Final positions: The value 3 is the page_id, and url identifies the page's URL.

This tuple describes a page with five incoming links, no earlier rank value available, a displayed new rank of 1.0, page ID 3, and the shown URL.

Connectivity Through Incoming Links

The first tuple element is the most direct measure of page connectivity. It counts how many other pages in the crawl link to the page represented by that tuple. A larger incoming-link count means that more pages in the crawled network point toward that page. Pages with many incoming links may be navigation hubs, homepages, or popular content pages, although the count alone does not determine the exact PageRank.

links tolinks tolinks toPage Alink sourcePage Blink sourcePage Clink sourceTarget pageincoming_links = 3
Which pages link to a target page, and how is that connectivity represented in the tuple?

The Output Filter

spdump.py does not print every page discovered by the spider. It displays only pages with at least one incoming link. A page with zero inbound connections is excluded by design, so absence from the output does not by itself prove that the spider failed to discover that page.

checkyesnoCrawled pagestored in crawl dataIncoming links ≥ 1Displayed tupleincluded in outputZero incoming linksexcluded from output
What determines whether a crawled page is included in the displayed results?

This filter keeps attention on pages that participate in the connectivity network. The source describes pages with no incoming links as leaf or terminal pages: they may have outgoing links, but no other page in the crawled network links back to them. Because they receive no links from other pages, they cannot receive PageRank from those pages.

  • Assuming that every discovered page must appear in spdump.py output.

    The tool filters out pages with no inbound connections.

    Fix: Interpret absence together with the output rule: only pages with at least one incoming link are displayed.

  • Treating a missing page as proof that the spider did not find it.

    The page may have been discovered but excluded because no crawled page links to it.

    Fix: Distinguish discovery from display. spdump.py shows a filtered connectivity summary, not an unfiltered list of every page.

Validation and Debugging

Use the tuples as evidence when checking a crawl. Compare the URLs shown with the pages you expected to discover. Inspect incoming-link counts to see whether an expected hub has meaningful connectivity. Check page IDs and rank fields for values that do not fit your understanding of the stored crawl. Unexpected links, surprisingly small link counts, missing URLs, or unexpected rank values can point to a crawl that did not go as deeply as expected, missed pages, or stored relationships different from the ones you intended to examine.

comparecompareunexpected resultunexpected resultExpected link countfor a central pageObserved link countin incoming_linksExpected URLpage to verifyDisplayed URLURL in tupleCrawl validationinvestigate differences
How can differences between what you expected and what spdump.py shows help identify crawl problems?

What do you think happens?

A page was discovered during the crawl, but no crawled page links to it. Will it appear in spdump.py output?

  • Yes, every discovered page is displayed
  • No, it is filtered out
  • Only if its page_id is small
Reveal answer

Answer: No, it is filtered out.

spdump.py displays only pages with at least one incoming link. A discovered page with zero inbound connections is excluded by design.

  • Verify that an expected URL is present when it should participate in the displayed connectivity network.
  • Read the first tuple element to check whether the observed incoming-link count matches the crawl structure you expected.
  • Use page_id and url together to identify which stored page a tuple describes.
  • Treat old_rank being None on the first run as an expected condition rather than automatically as an error.
  • Remember that missing output may result from the zero-incoming-link filter.

Practice Reading Tuples

EASY

Suppose a displayed tuple has the form (2, None, 0.5, 7, url). Explain what each of the five positions tells you, then state whether this page passed the spdump.py output filter and why.

Hints
  • Start with the first position because it determines whether the page has incoming connectivity.
  • Use the tuple order: incoming_links, old_rank, new_rank, page_id, url.
  • A page must have at least one incoming link to appear.

Practice answer

Interpret (2, None, 0.5, 7, url).

Connectivity: The page has two incoming links, meaning two other pages in the crawl link to it.

Ranks: old_rank is None, while new_rank is 0.5. None is consistent with an initial run before earlier numeric rank values are available.

Identity: The page_id is 7 and the final field is the page URL.

Filter decision: The page appears because its incoming-link count is two, which is at least one.

The tuple represents a displayed page with two incoming links, an unavailable earlier rank, a new rank of 0.5, page ID 7, and the shown URL.

Key Takeaways

  • spdump.py reads spider.sqlite and presents a human-readable connectivity summary.
  • Each output line follows the order (incoming_links, old_rank, new_rank, page_id, url).
  • incoming_links counts how many other crawled pages link to the page represented by the tuple.
  • Only pages with at least one incoming link appear; pages with zero inbound connections are filtered out.
  • Comparing URLs, link counts, IDs, and rank fields with your expectations helps validate and debug a crawl.