Introduction to SQLite
spider.py is the script used to crawl websites and record links.
A Crawl That Remembers
A web spider does more than visit one page. The spider.py script crawls websites, records links, and discovers new content by extracting links from the HTML of pages it visits. The spider.sqlite file provides persistent storage for the crawled pages and the connections between them.
The central relationship is simple: spider.py performs the crawl, while spider.sqlite stores the pages and connections found during that crawl.
Following the Crawl
The crawl begins with spider.py. As the script visits a page, it examines that page's HTML. Links found in the HTML become discoveries that can lead the spider toward additional pages. This link-extraction step is the mechanism that allows the crawl to find new content rather than stopping at the first page.
Tracing a Link Discovery
Imagine that spider.py visits a page whose HTML contains links to two other pages. What happens in the crawl?
Visit: spider.py visits the page as part of the website crawl.
Inspect: The spider examines the HTML of the visited page.
Extract: The spider discovers the links contained in that HTML.
Continue: The discovered links provide paths toward additional pages and new content.
The crawl advances because links extracted from a visited page identify further content to crawl.
What spider.sqlite Stores
spider.sqlite is the database where all crawled pages and their connections are stored. The database therefore records more than an isolated list of pages: it also stores how those pages are connected through discovered links.
Resetting the Crawl
To restart a crawl, delete the spider.sqlite file. Deleting this file removes the stored record of the crawled pages and their connections, which resets the persistent crawl state.
Assuming that changing spider.py alone resets the crawl.
The crawl information is stored in spider.sqlite, so the existing database file still contains the recorded pages and connections.
Fix:
Delete spider.sqlite when the goal is to restart the crawl.Treating spider.sqlite as the crawler itself.
spider.py is the script used to crawl websites and record links. spider.sqlite is where the crawled pages and their connections are stored.
Fix:
Assign crawling and link recording to spider.py, and assign persistent storage to spider.sqlite.Ignoring the HTML link-extraction step.
The spider discovers new content by extracting links from the HTML of visited pages.
Fix:
Trace the crawl through a visited page's HTML and the links found there.
Check Your Understanding
A spider visits a page and finds links in its HTML. Explain which component performs the crawl, which component stores the crawled pages and connections, and what must be deleted to restart the crawl.
Hints
- Separate the active script from the stored database file.
- Use the sequence visited page, HTML, extracted links, and additional content.
- The reset action concerns the file that stores crawl data.
What do you think happens?
If spider.sqlite is not deleted, has the crawl's stored state been reset?
Reveal answer
Answer: No
The stored crawl data remains in spider.sqlite. To restart the crawl, the spider.sqlite file must be deleted.
Key Takeaways
- spider.py is the script that crawls websites and records links.
- The spider discovers new content by extracting links from the HTML of visited pages.
- spider.sqlite stores all crawled pages and their connections.
- Deleting spider.sqlite resets the stored crawl state and allows the crawl to restart.
Key Takeaways
- A web spider advances by visiting pages and extracting links from their HTML.
- spider.py performs the crawling and records discovered links.
- spider.sqlite stores crawled pages and the connections between them.
- Deleting spider.sqlite is the required reset action for restarting the crawl.