Concepts / Building Network Applications

Building Network Applications

Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code, mimicking what a browser does but automating the data extraction process.

  • Programming

From Browsing to Automation

When you browse the web, you normally enter a URL, wait for a server to respond, and read the page that your browser renders. Web scraping changes the role of the reader: a program sends the request, receives the page's HTML, analyzes its structure, and extracts selected information. The program is performing a browser-like retrieval task, but it is automating the reading and data collection instead of displaying the page for a person.

Web scraping is a program-based technique for retrieving web pages and extracting specific data from their HTML source code. It mimics what a browser does while automating the data extraction process.

enters URLrequests pagerendered contentrequests pageHTML responseextractsPersonreads rendered pageBrowsersends requestWeb serverreturns page resourcesScraping programextracts selected dataWeb serverreturns HTMLExtracted datastored or processed
What is different about the actions and data flow when a person browses a page versus when a program scrapes it?

The Scraping Pipeline

The fundamental scraping workflow has five connected stages. First, the program sends an HTTP request for a web page. Next, it receives the page's HTML. It then parses the HTML structure so that the content can be examined, extracts the particular data it needs, and finally stores or processes the results. The important idea is that retrieval and extraction are separate activities: receiving HTML gives the program the source material, while parsing and extraction turn that source into useful information.

responseinspectselectstore or processHTTP requestask for a pageHTMLpage sourceParse structureanalyze HTMLTarget dataselected informationResultsstore or process
How does a scraping program request a web page, receive its HTML, and extract specific data from it?

A Focused Scraping Task

Suppose a program needs to collect headlines from a news feed.

Request: The program sends an HTTP request for the news page.

Receive: The server returns the page's HTML source.

Parse: The program examines the HTML structure to locate the information representing headlines.

Extract: The program selects the headlines rather than treating every part of the page as its target data.

Process: The collected headlines can be stored or processed as a dataset.

The program has completed a focused scraping task: it retrieved one page and extracted one type of information from it.

Following Links Across Pages

A scraper does not have to stop after processing one page. It can extract hyperlinks from the page's HTML and use those links to retrieve additional pages. Each newly retrieved page can provide more data and more links. This repeated process is web traversal: the program moves from a starting page through the relationships represented by hyperlinks.

extract link and requestextract link and requestextract link and requestPage Astarting pagePage Blinked pagePage Clinked pagePage Dnewly discovered page
How does a scraper follow links from one page to discover and process additional pages?

Imagine a program beginning with a directory page. It retrieves that page, extracts links to individual entries, retrieves those linked pages, and extracts information from each entry. The program is still using the same scraping pipeline on every page, but link extraction determines which page it processes next.

Link extraction turns a single-page data task into a traversal task. The scraper first analyzes the current page, then treats hyperlinks as paths to additional pages that can be retrieved and analyzed.

Search Engines at Scale

Search engines apply these principles on a much larger scale. Automated spiders retrieve pages, extract their content and hyperlinks, follow those links, and repeat the process. This systematic traversal helps search engines discover publicly accessible pages and build an index of the web.

retrieveextractextractcatalogdiscover more pagesmeasure frequencyWeb pagediscovered pageSpiderretrieves pagePage contentextracted informationHyperlinkspaths to pagesSearch indexcataloged pagesLink frequencyimportance signal
How do search engines retrieve pages, extract their content and links, and build an index of the web?
Web scrapingSearch engine spidering
Focused data extraction from web pagesLarge-scale traversal and indexing of the web
May visit a small number of pagesContinuously follows links to discover pages
Usually collects a specific datasetCatalogs content and builds an index
The general techniqueA specialized application of scraping

Search engines also analyze relationships between pages. The source material identifies link frequency as one signal of importance: when many pages link to a page, the search engine interprets that pattern as evidence that the page is valuable. Therefore, search-engine scraping is not limited to collecting page content; it also examines connections and patterns across pages.

Common Conceptual Mistakes

  • Treating scraping as the same thing as manually reading a web page.

    The two activities may involve the same web server, but their data flows and goals differ.

    Fix: Remember that scraping automates retrieval and extraction rather than asking a person to read the rendered page.

  • Stopping the explanation at the HTTP request.

    Retrieving HTML supplies the source material; it does not yet produce the selected dataset.

    Fix: Describe the complete sequence: request, receive HTML, parse, extract, and store or process.

  • Calling every scraper a search-engine spider.

    Scraping is focused extraction, while spidering is large-scale, continuous traversal used for indexing.

    Fix: Use scraper for the general extraction technique and spider for the specialized web-traversal application.

  • Ignoring hyperlinks after extracting page data.

    Following extracted links is the foundation of web traversal and is central to how search engines discover pages.

    Fix: Separate the current page's target data from its links to possible next pages.

Check Your Understanding

MEDIUM

Describe what happens when a scraping program starts with one web page and discovers two additional pages through hyperlinks. Your explanation should identify the request, the HTML response, parsing, data extraction, link extraction, and the reason the process can continue.

Hints
  • Start with the five stages of the single-page scraping workflow.
  • Distinguish the target data extracted from the links used for traversal.
  • Explain how the newly discovered pages become inputs to the same workflow.

What do you think happens?

A program retrieves a page's HTML but does not parse it or extract any target information. Has it completed the scraping workflow?

  • Yes, because receiving HTML is the entire task
  • No, because retrieval must be followed by parsing and extraction
  • Yes, because links are not part of scraping
  • No, because a person must read the page first
Reveal answer

Answer: No, because retrieval must be followed by parsing and extraction.

The fundamental workflow includes sending a request, receiving HTML, parsing its structure, extracting target data, and storing or processing the results.

Key Takeaways

  1. Web scraping uses a program to retrieve web pages and extract selected data from their HTML.
  2. The core workflow is request, receive HTML, parse, extract, and store or process.
  3. Link extraction lets a scraper move from one page to additional pages through web traversal.
  4. Search-engine spiders use these principles at scale to discover pages, build an index, and analyze link relationships.
  5. Scraping is a focused extraction technique, while spidering is a large-scale traversal strategy.

Key Takeaways

  • Web scraping is automated retrieval and extraction of data from web pages.
  • A scraper mimics the retrieval role of a browser but analyzes HTML instead of presenting a page for a person to read.
  • The scraping pipeline consists of requesting a page, receiving HTML, parsing it, extracting target data, and storing or processing the results.
  • Following extracted hyperlinks turns single-page scraping into web traversal.
  • Search-engine spiders apply scraping and link-following principles across the web to build an index and evaluate page relationships.