Concepts / Storing and Querying Spider Data with SQLite

Storing and Querying Spider Data with SQLite

spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).

  • Programming

From Database to Diagnostic View

After a spider finishes crawling, its findings are stored in spider.sqlite. That database file is a binary SQLite file, so it is not a convenient human-readable report. The spdump.py tool provides that report: it reads the database and displays information about the pages and connections discovered by the spider.

The most important habit is to read spdump.py as a diagnostic view rather than as a complete list of every page encountered. Its output focuses on pages that participate in the crawl's incoming-link network. That makes the output useful for checking connectivity, interpreting PageRank-related values, and spotting results that do not match your expectations.

readsdisplaysspider.sqlitestored crawl dataspdump.pyreads and formats dataTuple outputfive fields per displayedpage
How does crawl information move from the SQLite database into the displayed tuples?

Reading the Five Positions

Every displayed line follows this order: (incoming_links, old_rank, new_rank, page_id, url). Each position has a fixed meaning, so changing the order in your interpretation can lead to incorrect conclusions about the crawl.

position 0 to 1position 1 to 2position 2 to 3position 3 to 4incoming_linksnumber of linking pagesold_rankprevious rank valuenew_rankcurrent rank valuepage_idpage identifierurlpage address
What does each position in the spdump.py tuple represent?
  • incoming_links is the direct connectivity count: how many other pages in the crawl link to this page.
  • old_rank is the earlier rank value. It is often None on the first run.
  • new_rank is the newer rank value shown by the output after rank computation.
  • page_id is the page identifier stored in the tuple.
  • url is the page address associated with that record.

Interpreting One Tuple

Interpret the source example (5, None, 1.0, 3, url).

Position 1: The incoming-link count is 5, meaning five distinct pages in the crawled network contain a hyperlink pointing to this page.

Position 2: old_rank is None, which commonly occurs on the first run before multiple PageRank computations have produced earlier numeric values.

Position 3: new_rank is 1.0, the newer rank value shown for this record.

Positions 4 and 5: The page has page_id 3, and the final field supplies its URL.

This tuple describes a page with five incoming links, no earlier rank value available in the first-run situation, a displayed new rank of 1.0, page identifier 3, and the associated URL.

Tracking Rank Across Runs

The two rank fields let you notice whether rank information has an earlier value to compare with. On the first run, old_rank is often None. After PageRank has been computed multiple times, numeric values appear. Therefore, None in the old_rank position is not automatically evidence that the crawl failed; it can describe the state of an initial run.

after multiple computationsold_rankNoneold_ranknumeric valuenew_rankrank valuenew_rankrank value
How can the old_rank field change as PageRank is computed multiple times?

Understanding the Output Filter

spdump.py does not display every page discovered by the spider. It includes only pages with at least one incoming link. A page with zero inbound connections is excluded by design.

This filter focuses the report on pages that participate in the connectivity network. A page can be encountered in the crawl while still having no other crawled page link to it. Such a page has zero incoming links and therefore does not appear in spdump.py output.

inspect countyesnoCrawled pageIncoming linksat least one?Displayed tupleExcluded pagezero inbound connections
What determines whether a page from the SQLite data is included in the displayed results?

What do you think happens?

A page was discovered during the crawl but no other crawled page links to it. Will it appear in spdump.py output?

  • Yes, because every discovered page is displayed
  • No, because it has zero incoming links
Reveal answer

Answer: No, because it has zero incoming links.

spdump.py filters out pages with no inbound connections. Its output contains only pages with at least one incoming link.

Using Connectivity as Evidence

The first tuple position is especially useful when validating a crawl. If a page you expected to be central has only one incoming link, that result may indicate that the spider did not crawl deeply enough or missed pages. The count also helps you inspect the site's structure: pages with many incoming links can be navigation hubs, homepages, or popular content pages.

Incoming-link counts also provide the foundation for PageRank. Pages with more incoming links tend to have higher PageRank, although the exact calculation also depends on the PageRank of the pages linking to them. Use the rank fields together with the connectivity count instead of treating the count as the entire ranking calculation.

compareinvestigate mismatchCentral pagemany incoming linksexpectedCentral pageone incoming link observedCrawl validationinspect depth or missedpages
How can spdump.py output help compare an expected page network with the observed crawl?
  • Check whether important pages have the incoming-link counts you expected.
  • Use the URL and page_id fields to identify which record a surprising result belongs to.
  • Inspect old_rank and new_rank to understand whether the output reflects an initial run or later PageRank computation.
  • Remember that an absent page may have been filtered because its incoming-link count is zero.

Common Interpretation Mistakes

  • Treating spdump.py as a complete list of all discovered pages.

    The tool filters out pages with zero incoming links.

    Fix: Interpret the output as a filtered connectivity report, not as an unfiltered discovery list.

  • Reading the tuple fields in the wrong order.

    The format always begins with incoming_links and ends with url.

    Fix: Read every tuple using the fixed order: incoming_links, old_rank, new_rank, page_id, url.

  • Assuming old_rank being None proves that the spider failed.

    old_rank is often None on the first run.

    Fix: Recognize None as an expected initial-run state and inspect the rest of the tuple.

  • Treating incoming-link count as the complete PageRank calculation.

    The exact PageRank calculation also depends on the PageRank of the pages providing those links.

    Fix: Use incoming_links as a connectivity measure and examine rank fields separately.

Practice the Diagnostic Read

MEDIUM

Explain what you would investigate when an important page has a low incoming-link count, an old_rank value of None, or does not appear in spdump.py output.

Hints
  • Start with the first tuple position and identify what it measures.
  • Check whether the page may have zero incoming links and therefore be filtered.
  • Use the old_rank rule for an initial run before concluding that the result is an error.
  • Consider whether the spider crawled deeply enough or missed pages.

A Validation Checklist

Use spdump.py output to decide whether a crawl result deserves further investigation.

Identify the record: Use page_id and url to determine which page the tuple describes.

Check connectivity: Read incoming_links as the number of other crawled pages linking to that page.

Check visibility: If the page is missing from the output, consider whether its incoming-link count is zero.

Check rank state: Interpret old_rank as often being None on the first run and as numeric after multiple PageRank computations.

Compare with expectations: If an expected hub has very few incoming links, investigate crawl depth or missed pages.

The output becomes evidence for validating both the stored page network and the spider's behavior.

Key Takeaways

  1. spdump.py reads spider.sqlite and presents crawl information in a human-readable five-element tuple.
  2. The tuple order is incoming_links, old_rank, new_rank, page_id, url.
  3. incoming_links directly measures how many other crawled pages link to the page.
  4. Only pages with at least one incoming link appear in the output.
  5. old_rank is often None on the first run, while numeric values appear after PageRank has been computed multiple times.
  6. Unexpected counts, rank values, identifiers, URLs, or missing pages can help validate and debug the crawl.

Key Takeaways

  • spdump.py turns the spider.sqlite database into a readable view of the crawled page network.
  • Every displayed record uses the fixed format (incoming_links, old_rank, new_rank, page_id, url).
  • The incoming-link count shows direct page connectivity, while the rank fields show PageRank-related state.
  • Pages with zero incoming links are deliberately excluded from the output.
  • Comparing the output with expectations helps reveal shallow crawling, missed pages, and other crawl-validation issues.