Understanding PageRank Calculations
spdump.py reads spider.sqlite and displays pages in a five-element tuple format: (incoming_links, old_rank, new_rank, page_id, url).
From Crawl Data to Readable Evidence
After a web spider finishes crawling, it stores its collected data in spider.sqlite. That file is a binary SQLite database, so it is not a convenient human-readable view of the pages and connections the spider found. The spdump.py tool reads spider.sqlite and prints a summary that lets you inspect page connectivity and PageRank-related values.
The most useful way to approach spdump.py is as a validation window into the crawl. It does not merely list URLs. Each output line describes one page using five values: incoming_links, old_rank, new_rank, page_id, and url. Reading those values in the correct order helps you check whether the crawler discovered the relationships you expected.
Reading the Five Positions
Every line follows the same structure: (incoming_links, old_rank, new_rank, page_id, url). Position matters. The first value is not the page ID, and the final value is not a rank. Treat the tuple as an ordered record: each position answers a different question about the page.
| Position | Field | Meaning |
|---|---|---|
| 1 | incoming_links | The number of other pages in the crawl that link to this page. |
| 2 | old_rank | The earlier rank value; it is often None on the first run. |
| 3 | new_rank | The newer PageRank value shown by the tool. |
| 4 | page_id | The page's identifier in the stored crawl data. |
| 5 | url | The page address. |
The five fields printed for each page
Tracing One Tuple
Interpreting a sample output line
Interpret the tuple (5, None, 1.0, 3, url).
First position: The value 5 is incoming_links. Five distinct pages in the crawled network contain a hyperlink pointing to this page.
Second position: The value None is old_rank. The source describes this as common on the first run, before numeric rank values have appeared after multiple PageRank computations.
Third position: The value 1.0 is new_rank, the newer rank value displayed for the page.
Fourth position: The value 3 is page_id, the identifier associated with this page in the crawl data.
Fifth position: The value url identifies the page address represented by the tuple.
This tuple describes a page with five incoming links, no earlier numeric rank recorded in the displayed state, a new rank of 1.0, page ID 3, and the displayed URL.
The first value is the direct connectivity measure. It tells you how many other pages in the crawl link to the page represented by the rest of the tuple.
Why Some Pages Disappear
spdump.py does not print every page the spider discovered. Its output includes only pages with at least one incoming link. A page with zero inbound connections is filtered out by design, so its absence from the output does not by itself mean that the spider failed to discover it.
The filtering reflects a connectivity distinction. Some discovered pages may have outgoing links but no other pages in the crawled network linking to them. These are leaf or terminal pages in the network. Because they have zero incoming links, they cannot receive PageRank from other pages through inbound connections, so spdump.py keeps its displayed network focused on pages that participate in those connections.
If you expect a page to appear but cannot find it in spdump.py, first consider whether the page has any inbound links in the crawled network. Its absence may be the result of the output filter rather than a crawling failure.
Using Connectivity to Validate a Crawl
Use the first tuple value to compare the observed structure with your expectations. A page with many incoming links is highly connected within the crawled network. Pages with few incoming links are less connected. This information can reveal whether the spider seems to have captured the relationships you expected from the site.
- Check whether an expected central page has an incoming-link count that matches your expectations.
- If a page has fewer incoming links than expected, consider whether the spider crawled deeply enough or missed pages.
- Use page_id and url together to identify which stored page a tuple describes.
- Compare old_rank and new_rank to observe whether rank values have progressed beyond the initial state.
- Remember that a page absent from output may have zero inbound connections and therefore may have been filtered out.
Common Reading Mistakes
Treating the tuple as an unordered collection of values.
Each position has a fixed meaning. The first position is incoming_links, while page_id is the fourth position.
Fix:
Always map the tuple using the order incoming_links, old_rank, new_rank, page_id, url.Assuming an absent URL was never discovered.
spdump.py filters out pages with zero incoming links.
Fix:
Check whether the page may have zero inbound connections before diagnosing a crawl failure.Treating None in old_rank as a corrupted rank.
old_rank is often None on the first run; numeric values appear after PageRank has been computed multiple times.
Fix:
Interpret None in the context of the calculation stage.Using incoming_links as the complete PageRank explanation.
The exact PageRank calculation also depends on the PageRank of the pages linking to the page.
Fix:
Use incoming_links as a connectivity measure and inspect old_rank and new_rank as the rank fields.
Practice the Tuple Trace
Suppose an spdump.py line has the form (2, None, 1.0, 7, url). Without looking back at the table, identify what each of the five positions tells you. Then explain two reasons why a different crawled page might not appear in spdump.py output.
Hints
- Use the fixed field order rather than the data type of each value.
- Remember the output filter for pages with zero incoming links.
- Remember the special first-run meaning of old_rank being None.
Practice answer
Interpret (2, None, 1.0, 7, url) and explain why another page might be absent.
Connectivity: The first value, 2, means two other pages in the crawled network link to this page.
Earlier rank: The second value, None, is old_rank and is often seen on the first run.
Newer rank: The third value, 1.0, is new_rank.
Identity: The fourth value, 7, is page_id, and the fifth value is the page URL.
Filtering: Another crawled page may be absent because it has zero incoming links and spdump.py deliberately filters it out.
The tuple describes a page with two incoming links, an old_rank of None, a new_rank of 1.0, page ID 7, and the displayed URL. A page with no inbound connections would not be printed.
What to Remember
- spdump.py reads spider.sqlite and presents selected crawl data in a human-readable form.
- Every displayed record follows the order incoming_links, old_rank, new_rank, page_id, url.
- incoming_links directly measures how many other crawled pages link to the page.
- Only pages with at least one incoming link appear; pages with zero inbound connections are filtered out.
- old_rank is often None on the first run, while later numeric values show the progression of PageRank calculations.
- Use the output to validate page connectivity, investigate unexpected URLs or counts, and check whether the crawl matches your expectations.
Key Takeaways
- spdump.py provides a readable view of connectivity and rank-related data stored in spider.sqlite.
- The five tuple positions always mean incoming_links, old_rank, new_rank, page_id, and url.
- The incoming-link count shows how many other pages in the crawl link to the displayed page.
- Pages with zero incoming links are intentionally excluded from the output.
- Comparing the tuple fields with your crawl expectations helps debug and validate the spider's results.