Concepts / Web Crawling and Data Collection with spider.py

Web Crawling and Data Collection with spider.py

PageRank is computed by running sprank.py, which iteratively calculates importance scores based on the link structure in your database.

  • Programming

From Crawled Links to Rankings

A crawler collects pages and the links between them, but that collection does not yet tell you which pages are most important. PageRank supplies that next step. When you run sprank.py, it examines the link structure stored in the database and transforms it into numerical importance scores. The calculation is iterative: each pass uses the previous scores to produce a more refined set of scores.

addsprovides link structureproducesspider.pycrawl pagesPage databasepages and linkssprank.pyiterative calculationPageRank outputscores and records
What happens from collected pages to displayed PageRank results?

Starting an Iterative Run

To begin, invoke sprank.py and specify how many iterations you want. The program prompts you for that input, then performs the requested number of complete passes through the pages in the database. An iteration is not a single-page calculation. It is one pass through the collection, during which PageRank values are recalculated using the scores from the preceding iteration.

The iteration count controls how many refinement passes the program performs. Early scores are rough estimates; later passes use the updated scores to move the ranking toward a stable state.

A Two-Iteration Run

Trace a small PageRank calculation on a database containing five pages when sprank.py is run for two iterations.

Choose the run length: The program receives 2 as the requested number of iterations.

Complete iteration 1: The first output line represents one complete pass through all pages. Its average change is 0.547.

Complete iteration 2: The second pass uses the preceding scores. Its average change decreases to 0.227.

Read the trend: The decrease from 0.547 to 0.227 shows that the scores are changing less between passes.

The two iterations have started the ranking process, but the reported decrease should be understood as progress toward stabilization rather than proof that every possible later change is negligible.

containsinformsrefinesrefinesWeb pagesdatabase recordsInitial scoresrough estimatesIteration 1updated scoresIteration 2refined scoresLink structureincoming and outgoing links
How do page links flow into successive PageRank updates?

Reading Progress Output

Each progress line reports two pieces of information: the iteration number and the average change in PageRank score per page during that iteration. The average-change value is a diagnostic signal. A large value means the calculation is still making substantial corrections. A smaller value means the scores are changing less from one pass to the next.

change decreasesrefinement continueschange becomes very smallIteration 1average change 0.547Iteration 2average change 0.227Later passsmaller changeStable scoresnegligible change
How do score changes typically develop across iterations?

Interpreting Stored Scores

After the run, PageRank scores are stored in the database and can be queried. The scores are relative: they describe the importance of pages compared with other pages in the same network. A higher score indicates greater importance according to the link structure. The calculation reflects both how many pages point to a page and how important those referring pages are.

PagePageRank scoreInterpretation
Page 42.135Highest score in the stated simple output example
Page 20.659Tied with page 5 in the stated example
Page 50.659Tied with page 2 in the stated example
Page 10.559Slightly below pages 2 and 5 in the stated example

Illustrative PageRank output described in the source

A detailed database record adds context to the score. The source describes records containing a page ID, a version number of 1.0, the PageRank score, the number of outgoing links, and the URL. This lets you inspect both the ranking result and part of the page's link topology. For example, the source describes a record with a PageRank of 2.135 and four outgoing links.

ranks aboveties withranks abovePage 42.135Page 20.659Page 50.659Page 10.559
How do numerical PageRank values map to relative ordering?

Running the Cycle Again

PageRank is not necessarily a one-time operation. You can run sprank.py multiple times on the same database. A later run starts with the scores from the earlier run and iterates again, allowing further refinement when more convergence is needed.

The ranking workflow also changes when the crawler adds information. A practical cycle is to run spider.py, add newly crawled pages and their links to the database, and then run sprank.py again. The expanded network requires the scores to be recalculated so that the rankings account for the new pages and links.

ranked by sprank.pyexpanded by spider.pyranked againExisting databasepages and linksExisting rankingsprevious scoresExpanded databasenew pages and linksRecalculated rankingsupdated scores
How does a new crawl affect the next PageRank run?

Resetting Without Recrawling

Sometimes you want to restart the PageRank calculation without downloading the pages again. The spreset.py program resets all PageRank values to their initial state while preserving the page records and the link structure. After the reset, sprank.py can calculate scores again from the beginning using the same network.

starts overclears rank valuesready forRanked databasescores, pages, linksspreset.pyreset scoresInitial PageRankstatepages and links retainedsprank.pyrecalculate
What changes when PageRank values are reset, and what remains available?

Use a reset when you need a fresh PageRank calculation, want to test different parameters, or want to verify results while keeping the crawled pages and links. Do not treat reset as a replacement for crawling: it clears ranking values, not the stored network.

Common Diagnostic Mistakes

  • Treating the first iteration as the final ranking

    Initial estimates are rough, and later iterations refine them using the preceding scores.

    Fix: Inspect the average change across iterations and continue until the change is small enough for your acceptable convergence threshold.

  • Assuming that a larger score is an absolute measure of importance

    PageRank scores are relative to the other pages in the same network.

    Fix: Use the scores to compare pages within the same database and link structure.

  • Ignoring the average-change output

    The average change shows whether the calculation is still making substantial corrections or has begun to stabilize.

    Fix: Track whether the average change is decreasing and becoming very small.

  • Recrawling when only a fresh ranking calculation is needed

    spreset.py can reset rank values while preserving the crawled pages and link structure.

    Fix: Use spreset.py, then run sprank.py again when the network itself should remain unchanged.

Check Your Understanding

MEDIUM

You run sprank.py and observe average changes of 0.547 and then 0.227. The database has not changed since the run began. What does this pattern tell you, and what additional evidence would you look for before declaring convergence?

Hints
  • Compare the size of the second change with the first.
  • Convergence is associated with changes becoming very small.
  • Consider whether later iterations continue to reduce the average change below an acceptable threshold.

What do you think happens?

You need to restart PageRank but keep all crawled pages and links. Which action fits that goal?

  • Run spider.py again
  • Run spreset.py, then run sprank.py
  • Delete the page database
  • Use only the previous ranking output
Reveal answer

Answer: Run spreset.py, then run sprank.py.

spreset.py returns PageRank values to their initial state while preserving page records and link structure. sprank.py can then recalculate the rankings from that preserved network.

Operational Takeaways

  1. sprank.py converts the database's page-link structure into numerical PageRank scores through repeated passes.
  2. Each progress line reports an iteration number and the average change in score per page.
  3. A progressively smaller average change indicates that rankings are stabilizing; very small changes indicate convergence.
  4. Scores are relative to the pages in the same network, and detailed records can include IDs, scores, outgoing-link counts, and URLs.
  5. Run sprank.py again to refine or update rankings, use spider.py before ranking when new pages and links have been crawled, and use spreset.py when you need a fresh calculation without losing the network.

Key Takeaways

  • sprank.py iteratively calculates PageRank from the pages and links stored in the database.
  • The average change reported after each iteration reveals whether scores are still shifting or approaching stability.
  • PageRank values are relative rankings within the same network, and database records provide additional page and link metadata.
  • Repeated runs can refine rankings, while new spider.py results require ranking the expanded network again.
  • spreset.py clears PageRank values without removing the crawled pages or their links.