Analyzing Web Connectivity with PageRank
spider.py is the script used to crawl websites and record links.
From Pages to Connections
Analyzing web connectivity begins with collecting the connections between pages. In this learning path, spider.py is the script that crawls websites and records links. As it visits pages, it extracts links from their HTML and uses those discoveries to find additional content. The collected pages and their connections are stored in spider.sqlite.
The central mechanism is discovery through HTML: spider.py visits a page, extracts links from that page, and uses the discovered links to continue finding content.
Tracing One Crawl
A Small Crawl Trace
Suppose spider.py visits a page whose HTML contains links to two other pages. What information becomes available after that visit?
Visit: spider.py processes the visited page.
Inspect HTML: The script extracts links from the HTML of the visited page.
Record connections: The page and its connections are part of the crawl information recorded by spider.py.
Continue discovery: The extracted links provide additional content that the crawl can continue to discover.
One visited page has contributed its discovered links to the growing record of web connectivity.
This trace separates two related ideas. Visiting a page gives the spider HTML to inspect. Extracting links turns that HTML into connectivity information. The crawl is therefore not just a list of pages; it also records how those pages are connected.
What do you think happens?
If a visited page contains links, what is the next important discovery step?
Reveal answer
Answer: Extract the links from the page HTML
The source describes link extraction from visited HTML pages as the way the spider discovers new content.
Persistent Crawl Records
spider.sqlite is the database used to persist the crawl. It stores all crawled pages and their connections. This gives the crawl a durable record of both the pages that have been crawled and the relationships discovered between them.
| Component | Role |
|---|---|
| spider.py | Crawls websites and records links. |
| Visited page HTML | Provides links from which new content is discovered. |
| spider.sqlite | Stores crawled pages and their connections. |
The main parts of the crawl and their responsibilities.
Script and Database
The script and the database have different roles. spider.py performs the crawling work and records what it discovers. spider.sqlite is the persistent place where the resulting page and connection information is stored. Thinking of them as a pair helps explain why crawl information can remain available beyond the moment when a page is visited.
When reasoning about a crawl, ask two separate questions: what did spider.py discover, and where is that discovery stored? The first question concerns the script; the second concerns spider.sqlite.
Resetting the Crawl
A new crawl requires clearing the existing persistent crawl record. To restart the crawling process, delete the spider.sqlite file. Deleting this file removes the stored database that contains the crawled pages and their connections, so the crawl can begin again without that existing record.
Common Crawl Mistakes
Treating spider.py as the database
spider.py is the script used to crawl websites and record links, while spider.sqlite is the database where crawled pages and connections are stored.
Fix:
Keep the roles separate: spider.py performs the crawl, and spider.sqlite stores the crawl record.Ignoring page HTML when explaining discovery
The spider discovers new content by extracting links from the HTML of visited pages.
Fix:
Describe discovery as a link-extraction step performed on visited HTML.Trying to restart without deleting spider.sqlite
The required action for restarting a crawl is to delete spider.sqlite.
Fix:
Delete spider.sqlite before starting the crawl again.
Recording a Link
Imagine that a visited page contains one link to another page. The important transformation is from page HTML to extracted link, and then from extracted link to recorded connection. spider.py carries out the crawling and recording, while spider.sqlite preserves the resulting information as part of the crawl database.
Practice Check
A learner says: spider.py finds pages, but spider.sqlite only stores the page currently being visited. They also want to restart the crawl and plan to run spider.py again without changing any files. Identify both errors and state the correct explanation or action.
Hints
- Recall what spider.sqlite stores besides pages.
- Separate the crawler script from its persistent database.
- Recall the source's required action for restarting a crawl.
Practice Answer
Correct the learner's two errors.
Storage error: spider.sqlite stores all crawled pages and their connections, not only the page currently being visited.
Reset error: Running spider.py again without deleting spider.sqlite does not perform the specified reset action.
Correct action: Delete spider.sqlite to restart the crawl.
spider.py performs the crawl and records links; spider.sqlite stores the crawled pages and connections; deleting spider.sqlite resets the persistent crawl record.
Key Takeaways
- spider.py crawls websites and records links.
- The spider discovers new content by extracting links from the HTML of visited pages.
- spider.sqlite stores all crawled pages and their connections.
- The crawler script and its database have separate roles.
- To restart a crawl, delete the spider.sqlite file.
Key Takeaways
- A web spider builds connectivity information by extracting links from visited HTML pages.
- spider.py performs the crawl and records discovered links.
- spider.sqlite persistently stores crawled pages and their connections.
- Deleting spider.sqlite is the required way to restart the crawling process.