Concepts / Introduction to PageRank: Measuring Page Importance

Introduction to PageRank: Measuring Page Importance

spider.py systematically crawls web pages and stores them in a local database, recording both page content and the links between pages.

  • Programming

Why Collection Comes First

PageRank-style analysis cannot begin with an empty view of the web. Before you can study which pages are central or influential, you must first collect pages and the links between them. A web crawler creates this starting dataset by visiting pages, recording their content, and noting which pages link to which other pages.

The spider.py program is the collection engine in this process. It takes a starting URL, fetches the page, extracts its links, stores the page and its outgoing links in spider.sqlite, and then continues with another unvisited link. The result is a stored graph: pages are the graph's entities, and links are the relationships connecting them. Later, this graph can be queried, analyzed with algorithms such as page rank, or visualized with D3.

From Web Page to Stored Records

fetchstore pagerecord outgoing linkswritewriteLive web pageURL and page contentspider.pyfetches and extracts linksPage recordID, URL, outgoing-linkcountLink recordssource page ID and targetpage IDspider.sqlitepersistent local database
How does a page fetched from the web become stored page content and link records in the local database?

The crawler handles two kinds of information for each processed page. First, it stores a page record containing a unique page ID, the URL, and the number of outgoing links. Second, it stores relationships for those links. A relationship identifies the source page by its ID and the target page by its ID.

The database does not merely hold a list of URLs. It also preserves the direction of connections: one page is the source, and another page is the target.

Following the Unvisited Queue

beginpage arriveslinks foundcheck candidatesavailablenone available or limit reachedrepeatStarting URLFetch pageExtract linksStore page and linksUnvisited linksNext pageCrawl limit reached
What happens after the crawler reads links on one page, and how does it decide which pages to visit?

After fetching a page, spider.py extracts its links and adds the discovered relationships to the database. The crawler then selects another page that has not yet been visited. On a restart, it chooses a random unvisited page from its queue. The word systematic does not mean that every link is followed in a fixed order; it means that the crawler repeatedly applies the same discovery-and-storage process while respecting its crawl limit and its record of visited pages.

The Database Graph

containscontainscontainssource IDtarget IDPages tableID, URL, outgoing-linkcountPage Apage IDPage Bpage IDLinks tablesource page ID and targetpage IDPage A to Page Bsource to target
What contains what in the database, and how are stored pages connected to the links they contain?

The local SQLite file contains at least two tables: one for pages and one for links. The pages table represents the stored pages. The links table represents directed relationships between those pages. This separation allows later queries such as finding every page that links to a target, finding every page linked from a source, or calculating network statistics.

Suppose a stored page A links to pages B and C. The database keeps a page record for A, records that A has two outgoing links, and stores link relationships from A to B and from A to C. The graph can therefore preserve both the pages and the direction of their connections.

What a Restart Changes

writespersists between runsrecognizes existing recordsfetches unvisited pagesFirst crawl10 pages storedspider.sqlitestored pages and linksSecond crawlrequest 20 pagesStored pagesskippedNew pagesdatabase grows
How does a later crawling session recognize previously stored pages and continue without fetching them again unnecessarily?

The crawler is additive. Its database persists on disk, so a later run can use the records from an earlier run. If the first session collected 10 pages and a later session requests 20 pages, the crawler does not need to fetch the original 10 again. It skips pages already in the database and works toward the additional pages.

Tracing a Three-Page Crawl

Starting from example.com

Imagine an empty spider.sqlite database. You start spider.py at http://example.com/ and request three pages. Trace what is stored and how the crawler can continue.

Start with the requested URL: The crawler begins with http://example.com/ as its starting point and fetches that page.

Store the first page: The page receives a unique ID. Its URL and count of outgoing links are stored in the pages table. The links it contains are recorded as source-to-target relationships, and unvisited destinations become candidates for later crawling.

Choose another unvisited link: The crawler selects an unvisited link discovered during the process and fetches that page. It stores the new page record and its outgoing link records in the same database.

Continue until three pages are processed: The crawler repeats the same process for another unvisited page until the requested number of pages has been reached or no suitable unvisited page remains.

The database now contains page records and directed link records for the collected portion of the web. A later run can use this stored graph, skip pages already present, and continue from unvisited links.

Notice the distinction between discovery and storage. A link found in a page can be recorded as a relationship before its target page has been fetched and stored as a full page record. The crawler uses unvisited links as future work, gradually turning discovered connections into more complete page data.

Separate Starting Points

crawlscrawlsstores instores inStarting URL oneone webFirst page clusterpages and linksStarting URL twoanother webSecond page clusterpages and linksspider.sqliteone unified database
How can separate starting pages create independent crawl networks while sharing the same database?

A single database can contain more than one starting point. For example, one session can begin at http://www.dr-chuck.com/ and another can begin at http://www.wikipedia.org/. These separate starting points are called webs within the program. Their pages and links share spider.sqlite even if the two collections are not directly connected.

The crawler does not maintain completely separate queues for separate webs. It treats all unvisited links as one unified queue and may interleave visits from the different starting points.

Common Mistakes

  • Treating PageRank as the crawler itself.

    The crawler creates the graph data first. Page rank is an analysis that can be run after the database contains pages and their connections.

    Fix: Separate the workflow into collection first and analysis second.

  • Assuming that every discovered link has already been fetched.

    A discovered link can become an unvisited candidate before its target page is fetched.

    Fix: Distinguish between recording a relationship and processing the target page.

  • Expecting a second run to start from an empty database.

    The crawler is designed to remember pages already stored and avoid re-crawling them.

    Fix: Treat the existing database as accumulated crawl state. Delete spider.sqlite only when a fresh crawl is intended.

  • Assuming different starting URLs require different database files.

    Multiple webs can coexist in the same spider.sqlite database.

    Fix: Use the shared database while remembering that the starting points may remain unconnected.

  • Assuming the crawler always finishes one web before beginning another.

    All unvisited links are treated as one unified queue, so visits can be interleaved.

    Fix: Think of the database as one growing graph with potentially separate clusters.

Practice the Data Flow

MEDIUM

A database already contains pages collected from one starting URL. You run spider.py again with a different starting URL and request additional pages. Explain what happens to the existing page records, the new starting point, the unvisited queue, and the database structure.

Hints
  • Ask whether stored pages are fetched again.
  • Recall how the program treats multiple starting points.
  • Identify what the pages table and links table gain during the later session.

What do you think happens?

If two starting URLs are not directly connected, can their pages still be stored in the same spider.sqlite database?

  • Yes, as separate webs in one database
  • No, each starting URL requires its own database
  • Only if the crawler deletes the earlier pages
Reveal answer

Answer: Yes, as separate webs in one database

The crawler allows multiple independent starting points to coexist in the same database. It treats their unvisited links as one unified queue, even when the resulting webs are not directly connected.

Key Takeaways

  1. The crawler must collect pages and links before page-importance analysis can be performed.
  2. spider.py turns fetched pages into page records and directed link records in spider.sqlite.
  3. The database persists between runs, allowing later sessions to skip pages already stored and add new pages incrementally.
  4. Separate starting URLs can form separate webs while sharing one database.
  5. The resulting page-and-link graph can later be queried, analyzed with page rank, or visualized with D3.

Key Takeaways

  • A crawler provides the data foundation for analyzing page importance.
  • Each fetched page contributes both a page record and relationships to its outgoing links.
  • Persistent storage makes crawling additive rather than redundant.
  • Multiple independent starting points can share one database and one unified pool of unvisited links.
  • PageRank analysis becomes meaningful only after this graph of pages and links has been collected.