Concepts / Analyzing Web Connectivity with PageRank

Analyzing Web Connectivity with PageRank

spider.py is the script used to crawl websites and record links.

  • Programming

From Pages to Connections

Analyzing web connectivity begins with collecting the connections between pages. In this learning path, spider.py is the script that crawls websites and records links. As it visits pages, it extracts links from their HTML and uses those discoveries to find additional content. The collected pages and their connections are stored in spider.sqlite.

visitread HTMLfind linksmake availablevisit another pageStarting pagepage to visitVisited pageHTML contentExtract linkslinks in HTMLDiscovered contentlinks available forcrawlingContinue crawlrepeat with discoveredpages
How does spider.py discover a page, extract its links, make those links available for continued crawling, and repeat the process?

The central mechanism is discovery through HTML: spider.py visits a page, extracts links from that page, and uses the discovered links to continue finding content.

Tracing One Crawl

A Small Crawl Trace

Suppose spider.py visits a page whose HTML contains links to two other pages. What information becomes available after that visit?

Visit: spider.py processes the visited page.

Inspect HTML: The script extracts links from the HTML of the visited page.

Record connections: The page and its connections are part of the crawl information recorded by spider.py.

Continue discovery: The extracted links provide additional content that the crawl can continue to discover.

One visited page has contributed its discovered links to the growing record of web connectivity.

This trace separates two related ideas. Visiting a page gives the spider HTML to inspect. Extracting links turns that HTML into connectivity information. The crawl is therefore not just a list of pages; it also records how those pages are connected.

What do you think happens?

If a visited page contains links, what is the next important discovery step?

  • Extract the links from the page HTML
  • Delete spider.sqlite
  • Ignore the page connections
Reveal answer

Answer: Extract the links from the page HTML

The source describes link extraction from visited HTML pages as the way the spider discovers new content.

Persistent Crawl Records

spider.sqlite is the database used to persist the crawl. It stores all crawled pages and their connections. This gives the crawl a durable record of both the pages that have been crawled and the relationships discovered between them.

recordsrecordsstored instored inspider.pycrawl and record linksCrawled pagesstored page recordsPage connectionsstored link relationshipsspider.sqlitepersistent crawl database
What crawling data does spider.sqlite contain, and how does information move between spider.py and the database?
ComponentRole
spider.pyCrawls websites and records links.
Visited page HTMLProvides links from which new content is discovered.
spider.sqliteStores crawled pages and their connections.

The main parts of the crawl and their responsibilities.

Script and Database

The script and the database have different roles. spider.py performs the crawling work and records what it discovers. spider.sqlite is the persistent place where the resulting page and connection information is stored. Thinking of them as a pair helps explain why crawl information can remain available beyond the moment when a page is visited.

records crawl datacontainscontainsspider.pycrawlerspider.sqlitepersistent databaseCrawled pagesstored informationConnectionsstored relationships
How are the crawler script and its persistent database connected during a crawl?

When reasoning about a crawl, ask two separate questions: what did spider.py discover, and where is that discovery stored? The first question concerns the script; the second concerns spider.sqlite.

Resetting the Crawl

A new crawl requires clearing the existing persistent crawl record. To restart the crawling process, delete the spider.sqlite file. Deleting this file removes the stored database that contains the crawled pages and their connections, so the crawl can begin again without that existing record.

uses stored crawl datastarts recording againspider.pycrawl scriptspider.sqlitestored pages andconnectionsspider.pycrawl scriptNew crawl recordcreated as crawling startsagain
What changes when spider.sqlite is reset, and how does the crawler state differ before and after the reset?

Common Crawl Mistakes

  • Treating spider.py as the database

    spider.py is the script used to crawl websites and record links, while spider.sqlite is the database where crawled pages and connections are stored.

    Fix: Keep the roles separate: spider.py performs the crawl, and spider.sqlite stores the crawl record.

  • Ignoring page HTML when explaining discovery

    The spider discovers new content by extracting links from the HTML of visited pages.

    Fix: Describe discovery as a link-extraction step performed on visited HTML.

  • Trying to restart without deleting spider.sqlite

    The required action for restarting a crawl is to delete spider.sqlite.

    Fix: Delete spider.sqlite before starting the crawl again.

Recording a Link

extractrecordstoreVisited HTMLpage contentExtracted linkdiscovered connectionConnection recordlink recorded by spider.pyspider.sqlitepersistent storage
How does spider.py transform a link found on a webpage into crawl information stored for later use?

Imagine that a visited page contains one link to another page. The important transformation is from page HTML to extracted link, and then from extracted link to recorded connection. spider.py carries out the crawling and recording, while spider.sqlite preserves the resulting information as part of the crawl database.

Practice Check

MEDIUM

A learner says: spider.py finds pages, but spider.sqlite only stores the page currently being visited. They also want to restart the crawl and plan to run spider.py again without changing any files. Identify both errors and state the correct explanation or action.

Hints
  • Recall what spider.sqlite stores besides pages.
  • Separate the crawler script from its persistent database.
  • Recall the source's required action for restarting a crawl.

Practice Answer

Correct the learner's two errors.

Storage error: spider.sqlite stores all crawled pages and their connections, not only the page currently being visited.

Reset error: Running spider.py again without deleting spider.sqlite does not perform the specified reset action.

Correct action: Delete spider.sqlite to restart the crawl.

spider.py performs the crawl and records links; spider.sqlite stores the crawled pages and connections; deleting spider.sqlite resets the persistent crawl record.

Key Takeaways

  1. spider.py crawls websites and records links.
  2. The spider discovers new content by extracting links from the HTML of visited pages.
  3. spider.sqlite stores all crawled pages and their connections.
  4. The crawler script and its database have separate roles.
  5. To restart a crawl, delete the spider.sqlite file.

Key Takeaways

  • A web spider builds connectivity information by extracting links from visited HTML pages.
  • spider.py performs the crawl and records discovered links.
  • spider.sqlite persistently stores crawled pages and their connections.
  • Deleting spider.sqlite is the required way to restart the crawling process.