Web Crawling and Data Collection with spider.py
PageRank is computed by running sprank.py, which iteratively calculates importance scores based on the link structure in your database.
From Crawled Links to Rankings
A crawler collects pages and the links between them, but that collection does not yet tell you which pages are most important. PageRank supplies that next step. When you run sprank.py, it examines the link structure stored in the database and transforms it into numerical importance scores. The calculation is iterative: each pass uses the previous scores to produce a more refined set of scores.
Starting an Iterative Run
To begin, invoke sprank.py and specify how many iterations you want. The program prompts you for that input, then performs the requested number of complete passes through the pages in the database. An iteration is not a single-page calculation. It is one pass through the collection, during which PageRank values are recalculated using the scores from the preceding iteration.
The iteration count controls how many refinement passes the program performs. Early scores are rough estimates; later passes use the updated scores to move the ranking toward a stable state.
A Two-Iteration Run
Trace a small PageRank calculation on a database containing five pages when sprank.py is run for two iterations.
Choose the run length: The program receives 2 as the requested number of iterations.
Complete iteration 1: The first output line represents one complete pass through all pages. Its average change is 0.547.
Complete iteration 2: The second pass uses the preceding scores. Its average change decreases to 0.227.
Read the trend: The decrease from 0.547 to 0.227 shows that the scores are changing less between passes.
The two iterations have started the ranking process, but the reported decrease should be understood as progress toward stabilization rather than proof that every possible later change is negligible.
Reading Progress Output
Each progress line reports two pieces of information: the iteration number and the average change in PageRank score per page during that iteration. The average-change value is a diagnostic signal. A large value means the calculation is still making substantial corrections. A smaller value means the scores are changing less from one pass to the next.
Interpreting Stored Scores
After the run, PageRank scores are stored in the database and can be queried. The scores are relative: they describe the importance of pages compared with other pages in the same network. A higher score indicates greater importance according to the link structure. The calculation reflects both how many pages point to a page and how important those referring pages are.
| Page | PageRank score | Interpretation |
|---|---|---|
| Page 4 | 2.135 | Highest score in the stated simple output example |
| Page 2 | 0.659 | Tied with page 5 in the stated example |
| Page 5 | 0.659 | Tied with page 2 in the stated example |
| Page 1 | 0.559 | Slightly below pages 2 and 5 in the stated example |
Illustrative PageRank output described in the source
A detailed database record adds context to the score. The source describes records containing a page ID, a version number of 1.0, the PageRank score, the number of outgoing links, and the URL. This lets you inspect both the ranking result and part of the page's link topology. For example, the source describes a record with a PageRank of 2.135 and four outgoing links.
Running the Cycle Again
PageRank is not necessarily a one-time operation. You can run sprank.py multiple times on the same database. A later run starts with the scores from the earlier run and iterates again, allowing further refinement when more convergence is needed.
The ranking workflow also changes when the crawler adds information. A practical cycle is to run spider.py, add newly crawled pages and their links to the database, and then run sprank.py again. The expanded network requires the scores to be recalculated so that the rankings account for the new pages and links.
Resetting Without Recrawling
Sometimes you want to restart the PageRank calculation without downloading the pages again. The spreset.py program resets all PageRank values to their initial state while preserving the page records and the link structure. After the reset, sprank.py can calculate scores again from the beginning using the same network.
Use a reset when you need a fresh PageRank calculation, want to test different parameters, or want to verify results while keeping the crawled pages and links. Do not treat reset as a replacement for crawling: it clears ranking values, not the stored network.
Common Diagnostic Mistakes
Treating the first iteration as the final ranking
Initial estimates are rough, and later iterations refine them using the preceding scores.
Fix:
Inspect the average change across iterations and continue until the change is small enough for your acceptable convergence threshold.Assuming that a larger score is an absolute measure of importance
PageRank scores are relative to the other pages in the same network.
Fix:
Use the scores to compare pages within the same database and link structure.Ignoring the average-change output
The average change shows whether the calculation is still making substantial corrections or has begun to stabilize.
Fix:
Track whether the average change is decreasing and becoming very small.Recrawling when only a fresh ranking calculation is needed
spreset.py can reset rank values while preserving the crawled pages and link structure.
Fix:
Use spreset.py, then run sprank.py again when the network itself should remain unchanged.
Check Your Understanding
You run sprank.py and observe average changes of 0.547 and then 0.227. The database has not changed since the run began. What does this pattern tell you, and what additional evidence would you look for before declaring convergence?
Hints
- Compare the size of the second change with the first.
- Convergence is associated with changes becoming very small.
- Consider whether later iterations continue to reduce the average change below an acceptable threshold.
What do you think happens?
You need to restart PageRank but keep all crawled pages and links. Which action fits that goal?
Reveal answer
Answer: Run spreset.py, then run sprank.py.
spreset.py returns PageRank values to their initial state while preserving page records and link structure. sprank.py can then recalculate the rankings from that preserved network.
Operational Takeaways
- sprank.py converts the database's page-link structure into numerical PageRank scores through repeated passes.
- Each progress line reports an iteration number and the average change in score per page.
- A progressively smaller average change indicates that rankings are stabilizing; very small changes indicate convergence.
- Scores are relative to the pages in the same network, and detailed records can include IDs, scores, outgoing-link counts, and URLs.
- Run sprank.py again to refine or update rankings, use spider.py before ranking when new pages and links have been crawled, and use spreset.py when you need a fresh calculation without losing the network.
Key Takeaways
- sprank.py iteratively calculates PageRank from the pages and links stored in the database.
- The average change reported after each iteration reveals whether scores are still shifting or approaching stability.
- PageRank values are relative rankings within the same network, and database records provide additional page and link metadata.
- Repeated runs can refine rankings, while new spider.py results require ranking the expanded network again.
- spreset.py clears PageRank values without removing the crawled pages or their links.