Building Network Applications
Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code, mimicking what a browser does but automating the data extraction process.
From Browsing to Automation
When you browse the web, you normally enter a URL, wait for a server to respond, and read the page that your browser renders. Web scraping changes the role of the reader: a program sends the request, receives the page's HTML, analyzes its structure, and extracts selected information. The program is performing a browser-like retrieval task, but it is automating the reading and data collection instead of displaying the page for a person.
Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code. It mimics what a browser does while automating the data extraction process.
The Scraping Pipeline
The fundamental scraping workflow has five connected stages. First, the program sends an HTTP request for a web page. Next, it receives the page's HTML. It then parses the HTML structure so that the content can be examined, extracts the particular data it needs, and finally stores or processes the results. The important idea is that retrieval and extraction are separate activities: receiving HTML gives the program the source material, while parsing and extraction turn that source into useful information.
A Focused Scraping Task
Suppose a program needs to collect headlines from a news feed.
Request: The program sends an HTTP request for the news page.
Receive: The server returns the page's HTML source.
Parse: The program examines the HTML structure to locate the information representing headlines.
Extract: The program selects the headlines rather than treating every part of the page as its target data.
Process: The collected headlines can be stored or processed as a dataset.
The program has completed a focused scraping task: it retrieved one page and extracted one type of information from it.
Following Links Across Pages
A scraper does not have to stop after processing one page. It can extract hyperlinks from the page's HTML and use those links to retrieve additional pages. Each newly retrieved page can provide more data and more links. This repeated process is web traversal: the program moves from a starting page through the relationships represented by hyperlinks.
Imagine a program beginning with a directory page. It retrieves that page, extracts links to individual entries, retrieves those linked pages, and extracts information from each entry. The program is still using the same scraping pipeline on every page, but link extraction determines which page it processes next.
Link extraction turns a single-page data task into a traversal task. The scraper first analyzes the current page, then treats hyperlinks as paths to additional pages that can be retrieved and analyzed.
Search Engines at Scale
Search engines apply these principles on a much larger scale. Automated spiders retrieve pages, extract their content and hyperlinks, follow those links, and repeat the process. This systematic traversal helps search engines discover publicly accessible pages and build an index of the web.
| Web scraping | Search engine spidering |
|---|---|
| Focused data extraction from web pages | Large-scale traversal and indexing of the web |
| May visit a small number of pages | Continuously follows links to discover pages |
| Usually collects a specific dataset | Catalogs content and builds an index |
| The general technique | A specialized application of scraping |
Search engines also analyze relationships between pages. The source material identifies link frequency as one signal of importance: when many pages link to a page, the search engine interprets that pattern as evidence that the page is valuable. Therefore, search-engine scraping is not limited to collecting page content; it also examines connections and patterns across pages.
Common Conceptual Mistakes
Treating scraping as the same thing as manually reading a web page.
The two activities may involve the same web server, but their data flows and goals differ.
Fix:
Remember that scraping automates retrieval and extraction rather than asking a person to read the rendered page.Stopping the explanation at the HTTP request.
Retrieving HTML supplies the source material; it does not yet produce the selected dataset.
Fix:
Describe the complete sequence: request, receive HTML, parse, extract, and store or process.Calling every scraper a search-engine spider.
Scraping is focused extraction, while spidering is large-scale, continuous traversal used for indexing.
Fix:
Use scraper for the general extraction technique and spider for the specialized web-traversal application.Ignoring hyperlinks after extracting page data.
Following extracted links is the foundation of web traversal and is central to how search engines discover pages.
Fix:
Separate the current page's target data from its links to possible next pages.
Check Your Understanding
Describe what happens when a scraping program starts with one web page and discovers two additional pages through hyperlinks. Your explanation should identify the request, the HTML response, parsing, data extraction, link extraction, and the reason the process can continue.
Hints
- Start with the five stages of the single-page scraping workflow.
- Distinguish the target data extracted from the links used for traversal.
- Explain how the newly discovered pages become inputs to the same workflow.
What do you think happens?
A program retrieves a page's HTML but does not parse it or extract any target information. Has it completed the scraping workflow?
Reveal answer
Answer: No, because retrieval must be followed by parsing and extraction.
The fundamental workflow includes sending a request, receiving HTML, parsing its structure, extracting target data, and storing or processing the results.
Key Takeaways
- Web scraping uses a program to retrieve web pages and extract selected data from their HTML.
- The core workflow is request, receive HTML, parse, extract, and store or process.
- Link extraction lets a scraper move from one page to additional pages through web traversal.
- Search-engine spiders use these principles at scale to discover pages, build an index, and analyze link relationships.
- Scraping is a focused extraction technique, while spidering is a large-scale traversal strategy.
Key Takeaways
- Web scraping is automated retrieval and extraction of data from web pages.
- A scraper mimics the retrieval role of a browser but analyzes HTML instead of presenting a page for a person to read.
- The scraping pipeline consists of requesting a page, receiving HTML, parsing it, extracting target data, and storing or processing the results.
- Following extracted hyperlinks turns single-page scraping into web traversal.
- Search-engine spiders apply scraping and link-following principles across the web to build an index and evaluate page relationships.